This is a general question about the procedures concerning text mining. Suppose one has a Corpus of documents classified as Spam/No_Spam. As standard procedure one pre-process the data, removing punctuation, stops words etc. After converting it into a DocumentTermMatrix one can build some models to predict spam/No_Spam. Here is my problem. Now I want to use the model built for new documents arrive. In order to check a single document I would have to build a DocumentTerm*Vector*? so it can be used to predict Spam/No_Spam. In the documentation of tm I found one converts the full Corpus into a Matrix using for example tfidf weights. How can I then convert a single vector using the idf from the Corpus? do i have to change my corpus and build a new DocumentTermMatrix every time? I processed my corpus, converted it into a matrix and then split it into a Training and Testing sets. But here the test set was built in the same line as the document matrix of the full set. I can check precision etc, but do not know whats the best procedure for new text classification.
Ben, Imagine I have a preprocessed DocumentTextMatrix, I convert it into a data.frame.
dtm <- DocumentTermMatrix(CorpusProc,control = list(weighting =function(x) weightTfIdf(x, normalize =FALSE),stopwords = TRUE, wordLengths=c(3, Inf), bounds = list(global = c(4,Inf))))
dtmDataFrame <- as.data.frame(inspect(dtm))
Added a Factor Variable and built a model.
Corpus.svm<-svm(Risk_Category~.,data=dtmDataFrame)
Now imagine I give you a new document d (was not in your Corpus before) and you want to know the model prediction spam/No_Spam. How you do that?
Ok lets create an example based on the code used here.
examp1 <- "When discussing performance with colleagues, teaching, sending a bug report or searching for guidance on mailing lists and here on SO, a reproducible example is often asked and a