From the course: Natural Language Processing in Python

Unlock this course with a free trial

Join today to access over 26,400 courses taught by industry experts.

CountVectorizer in Python

CountVectorizer in Python

“

Within Python, you would create a count vectorizer object to make a document term matrix. Here's what the code for that would look like. You can see we're first importing count vectorizer from scikit-learn, and in typical scikit-learn fashion, the first thing we have to do is create a count vectorizer object. Now within this count vectorizer object, we're specifying three parameters here, stop words, n-gram range, and mindf. These are all optional parameters, but they're the three that I use most frequently. So first, stop words. This is a concept you're already familiar with. These are words that don't have much meaning and within your count vectorizer object you can specify a language here. So you can say remove all English stop words. By default, no stop words are removed. There's also this n-gram range and this is where you set your term length. Do you want to see unigrams or one words, bigrams, two words, trigrams, three words and so on. So by default the n-gram range is 1 to 1…

Contents