Sentiment analysis with machine learning in R

Machine learning makes sentiment analysis more convenient. This post would introduce how to do sentiment analysis with machine learning using R. In the landscape of R, the sentiment R package and the more general text mining package have been well developed by Timothy P. Jurka. You can check out the sentiment package and the fantastic RTextTools package. Actually, Timothy also writes an maxent package for low-memory multinomial logistic regression (also known as maximum entropy).

However, the naive bayes method is not included into RTextTools. The e1071 package did a good job of implementing the naive bayes method. e1071 is a course of the Department of Statistics (e1071), TU Wien. Its primary developer is David Meyer.

It is still necessary to learn more about text analysis. Text analysis in R has been well recognized (see the R views on natural language processing). Part of the success belongs to the tm package: A framework for text mining applications within R. It did a good job for text cleaning (stemming, delete the stopwords, etc) and transforming texts to document-term matrix (dtm). There is one paper about it. As you know the most important part of text analysis is to get the feature vectors for each document. The word feature is the most important one. Of course, you can also extend the unigram word features to bigram and trigram, and so on to n-grams. However, here for our simple case, we stick to the unigram word features.

Note that it’s easy to use ngrams in R. In the past, the package of Rweka supplies functions to do it, check this example. Now, you can set the ngramLength in the function of create_matrix using RTextTools.

The first step is to read data:

library(RTextTools)
library(e1071)

pos_tweets =  rbind(
  c('I love this car', 'positive'),
  c('This view is amazing', 'positive'),
  c('I feel great this morning', 'positive'),
  c('I am so excited about the concert', 'positive'),
  c('He is my best friend', 'positive')
)

neg_tweets = rbind(
  c('I do not like this car', 'negative'),
  c('This view is horrible', 'negative'),
  c('I feel tired this morning', 'negative'),
  c('I am not looking forward to the concert', 'negative'),
  c('He is my enemy', 'negative')
)

test_tweets = rbind(
  c('feel happy this morning', 'positive'),
  c('larry friend', 'positive'),
  c('not like that man', 'negative'),
  c('house not great', 'negative'),
  c('your song annoying', 'negative')
)

tweets = rbind(pos_tweets, neg_tweets, test_tweets)

Then we can build the document-term matrix:

# build dtm
matrix= create_matrix(tweets[,1], language="english", 
                      removeStopwords=FALSE, removeNumbers=TRUE, 
                      stemWords=FALSE) 

Now, we can train the naive Bayes model with the training set. Note that, e1071 asks the response variable to be numeric or factor. Thus, we convert characters to factors here. This is a little trick.

# train the model
mat = as.matrix(matrix)
classifier = naiveBayes(mat[1:10,], as.factor(tweets[1:10,2]) )

Now we can step further to test the accuracy.

# test the validity
predicted = predict(classifier, mat[11:15,]); predicted
table(tweets[11:15, 2], predicted)
recall_accuracy(tweets[11:15, 2], predicted)

Apparently, the result is the same with Python (compare it with the results in an another post).

How about the other machine learning methods?

As I mentioned, we can do it using RTextTools. Let’s rock!

First, to specify our data:

# build the data to specify response variable, training set, testing set.
container = create_container(matrix, as.numeric(as.factor(tweets[,2])),
                             trainSize=1:10, testSize=11:15,virgin=FALSE)

Second, to train the model with multiple machine learning algorithms:

models = train_models(container, algorithms=c("MAXENT" , "SVM", "RF", "BAGGING", "TREE"))

Now, we can classify the testing set using the trained models.

results = classify_models(container, models)

How about the accuracy?

# accuracy table
table(as.numeric(as.factor(tweets[11:15, 2])), results[,"FORESTS_LABEL"])
table(as.numeric(as.factor(tweets[11:15, 2])), results[,"MAXENTROPY_LABEL"])

# recall accuracy
recall_accuracy(as.numeric(as.factor(tweets[11:15, 2])), results[,"FORESTS_LABEL"])
recall_accuracy(as.numeric(as.factor(tweets[11:15, 2])), results[,"MAXENTROPY_LABEL"])
recall_accuracy(as.numeric(as.factor(tweets[11:15, 2])), results[,"TREE_LABEL"])
recall_accuracy(as.numeric(as.factor(tweets[11:15, 2])), results[,"BAGGING_LABEL"])
recall_accuracy(as.numeric(as.factor(tweets[11:15, 2])), results[,"SVM_LABEL"])

To summarize the results (especially the validity) in a formal way:

# model summary
analytics = create_analytics(container, results)
summary(analytics)
head(analytics@document_summary)
analytics@ensemble_summar

To cross validate the results:

N=4
set.seed(2014)
cross_validate(container,N,"MAXENT")
cross_validate(container,N,"TREE")
cross_validate(container,N,"SVM")
cross_validate(container,N,"RF")

The results can be found on my Rpub page. It seems that maxent reached the same recall accuracy as naive Bayes. The other methods even did a worse job. This is understandable, since we have only a very small data set. To enlarge the training set, we can get a much better results for sentiment analysis of tweets using more sophisticated methods. I will show the results with anther example.

Sentiment analysis for tweets

The data comes from victorneo. victorneo shows how to do sentiment analysis for tweets using Python. Here, I will demonstrate how to do it in R.

Read data:

###################
"load data"
###################
setwd("D:/Twitter-Sentimental-Analysis-master/")
happy = readLines("./happy.txt")
sad = readLines("./sad.txt")
happy_test = readLines("./happy_test.txt")
sad_test = readLines("./sad_test.txt")

tweet = c(happy, sad)
tweet_test= c(happy_test, sad_test)
tweet_all = c(tweet, tweet_test)
sentiment = c(rep("happy", length(happy) ), 
              rep("sad", length(sad)))
sentiment_test = c(rep("happy", length(happy_test) ), 
                   rep("sad", length(sad_test)))
sentiment_all = as.factor(c(sentiment, sentiment_test))

library(RTextTools)

First, try naive Bayes.

# naive bayes
mat= create_matrix(tweet_all, language="english", 
                   removeStopwords=FALSE, removeNumbers=TRUE, 
                   stemWords=FALSE, tm::weightTfIdf)

mat = as.matrix(mat)

classifier = naiveBayes(mat[1:160,], as.factor(sentiment_all[1:160]))
predicted = predict(classifier, mat[161:180,]); predicted

table(sentiment_test, predicted)
recall_accuracy(sentiment_test, predicted)

Then, try the other methods:

# the other methods
mat= create_matrix(tweet_all, language="english", 
                   removeStopwords=FALSE, removeNumbers=TRUE, 
                   stemWords=FALSE, tm::weightTfIdf)

container = create_container(mat, as.numeric(sentiment_all),
                             trainSize=1:160, testSize=161:180,virgin=FALSE) #可以设置removeSparseTerms

models = train_models(container, algorithms=c("MAXENT",
                                              "SVM",
                                              #"GLMNET", "BOOSTING", 
                                              "SLDA","BAGGING", 
                                              "RF", # "NNET", 
                                              "TREE" 
))

# test the model
results = classify_models(container, models)
table(as.numeric(as.numeric(sentiment_all[161:180])), results[,"FORESTS_LABEL"])
recall_accuracy(as.numeric(as.numeric(sentiment_all[161:180])), results[,"FORESTS_LABEL"])

Here we also want to get the formal test results, including:

  • analytics@algorithm_summary: Summary of precision, recall, f-scores, and accuracy sorted by topic code for each algorithm
  • analytics@label_summary: Summary of label (e.g. Topic) accuracy
  • analytics@document_summary: Raw summary of all data and scoring
  • analytics@ensemble_summary: Summary of ensemble precision/coverage. Uses the n variable passed into create_analytics()

Now let’s see the results:

# formal tests
analytics = create_analytics(container, results)
summary(analytics)

head(analytics@algorithm_summary)
head(analytics@label_summary)
head(analytics@document_summary)
analytics@ensemble_summary # Ensemble Agreement

# Cross Validation
N=3
cross_SVM = cross_validate(container,N,"SVM")
cross_GLMNET = cross_validate(container,N,"GLMNET")
cross_MAXENT = cross_validate(container,N,"MAXENT")

You can find that compared with naive Bayes, the other algorithms did a much better job to achieve a recall accuracy higher than 0.95. Check the results on Rpub.

If you have any question, feel free to post a comment below.

23 Comments

  1. PC
    Priyavrat Chauhan April 3, 2021

    Please give example, How can we use different feature extraction methods (BoW, TF_IDF and word2vector) in the above code. Thanks..

    Reply
  2. MJ
    Ms Jabeen January 25, 2019

    I am getting this error

    Error in create_matrix(tweets[, 1], language = “english”, removeStopwords = FALSE, :
    could not find function “create_matrix”

    Reply
    1. MJ
      Ms Jabeen January 25, 2019

      resolved. Thanks

      Reply
  3. DA
    Damian Anyamele June 25, 2018

    Thank you Wang!
    Please I am looking for a source code to do Aspect Base Sentiment Analysis in R programming. Could Anyone please help?

    Reply
  4. JP
    Júlio Piubello April 23, 2018

    Awesome, Awesome, Awesome, Awesome, Awesome 1000x awesome !!

    Reply
  5. D
    Deepan March 26, 2018

    I have done sentiment analysis and getting an accuracy of 0.8. Can you tell me how to get accuracy of 1

    Reply
  6. D
    Deepan March 22, 2018

    Hi,
    I have done some sentiment analysis but am getting an error saying subscript out of bonds.
    The code looks something like this:

    > library(RTextTools)
    > library(e1071)
    > library(SparseM)
    > pos_feeds = rbind(
    +
    + c(‘Stock market is positive’,’positive’),
    + c(‘GDP numbers are good’,’positive’),
    + c(‘Nifty is positive’,’positive’),
    + c(‘CPI numbers are good’,’positive’)
    + )
    > neg_feeds = rbind(
    +
    + c(‘Economy slow down’,’negative’),
    + c(‘Banking stocks plunge’,’negative’),
    + c(‘RBI hike repo rate’,’negative’),
    + c(‘Markets crash’,’negative’),
    + c(‘Nifty closes below its 200 DMA’,’negative’)
    + )
    > test_feeds = rbind(
    +
    + c(‘fGlobal cues are good’,’positive’),
    + c(‘US market ends positive’,’positive’),
    + c(‘Brexit’,’negative’),
    + c(‘PIIGS economy slows down’,’negative’),
    + c(‘Fed raise rates’,’negative’)
    + )
    >
    > feeds = rbind(pos_feeds,neg_feeds,test_feeds)
    >
    > # build dtm
    > matrix = create_matrix(feeds[,1],language = “english”,removeStopwords = FALSE,removeNumbers = TRUE,stemWords = FALSE)
    >
    > #train the model
    > mat = as.matrix(matrix)
    > classifier = naiveBayes(mat[1:10,], as.factor(feeds[1:10,2]) )
    >
    > # test the validity
    > predicted = predict(classifier, mat[11:15,]); predicted
    Error in mat[11:15, ] : subscript out of bounds
    > table(feeds[11:15, 2], predicted)
    Error in feeds[11:15, 2] : subscript out of bounds
    > recall_accuracy(feeds[11:15, 2], predicted)
    Error in feeds[11:15, 2] : subscript out of bounds

    Reply
  7. RL
    Renato Falcon Lyke July 7, 2017

    Hi,
    I have done some sentiment analysis for feedback we receive. I am trying to tie the feedback to a particular direction. Like whether we have received it from North South etc. I am unable to do this. The data is available in a CSV file with two columns Direction, Comments

    Any suggestions how could i go about it.

    Regards,
    Ren.

    Reply
  8. S
    shina145 May 26, 2017

    hello, i have done the sentiment analysis for some airline reviews … but i am facing a problem of interpreting the results of naive Bayes, SVM and the others . so my question is how do i interpret these results? (how do i make statistical sense out of it for my final year research? ) Thank You!

    Reply
  9. FM
    Fatima Mirza May 10, 2017

    Hello, how do I extend this code to add another emotion? Say for example, anger? Thank you!

    Reply
  10. FM
    Fatima Mirza May 9, 2017

    Hello,
    How do I use another dataset, not twitter dataset that is? Thank you!

    Reply
  11. K
    kiru April 22, 2017

    hi…. is it possible to show the output in chart representation???

    Reply
  12. D
    Davod March 1, 2017

    Hello … what a nice job !!! I’d like to do the same buyback on Facebook comments/private messages….any clue about how to sort out the text data to be run in R? Thank you very much

    Reply
  13. SK
    Shweta Kalla January 18, 2017

    Hi,
    The output matrix generated using the create_matrix function above generates a matrix with weights as tf and not tf-idf as it should have basis the below code (Copied from above). Can you plz explain what can change that.

    mat= create_matrix(tweet_all, language=”english”,

    removeStopwords=FALSE, removeNumbers=TRUE,

    stemWords=FALSE, tm::weightTfIdf)

    Thanks!

    Reply
  14. KT
    Ketan Thakare May 10, 2016

    I had done the sentiment analysis on twitter data using same metho,but i got low cross validation accuracy, so tell me how to increase accuracy using twitter dataset

    Reply
  15. KT
    Ketan Thakare May 10, 2016

    i have try for sentiment analysis of tweets using same method but i got low cross validation accuracy ..its only between 45 to 50 %.How i can increase accuracy for my dataset
    L

    Reply
    1. RA
      REBEEN ALI June 2, 2016

      hiii katena Thakare please can we talk about your project can I have your email please

      Reply
      1. KT
        Ketan Thakare June 11, 2016

        hi REBBEN my mail id is [email protected]

        Reply
  16. SA
    Solomon AathiRaj March 9, 2016

    What’s the point in classifying the sentiments having collected the happy & sad tweets itself? ML algorithms should classify it right?

    Reply
    1. D
      Diego March 15, 2016

      I think this method is supervised machine learning, you need to input some correct data and then the ML algorithm will learn and classify the future tweets

      Reply
  17. TM
    Tushar Mehrotra March 7, 2016

    could plz help me with some code on how to scrape twitter in order to fetch the posts and then classify them as positive, negative or neutral

    Reply
  18. TM
    Tushar Mehrotra March 7, 2016

    Hey im getting an error message when running > table(sentiment_test, predicted)….all arg must be of same length

    Reply
  19. S
    Sandy January 12, 2016

    super, very good, thanks

    Reply

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.