Week #49 – Stemming

In language processing, stemming is the process of taking multiple forms of the same word and reducing them to the same basic core form.

Comments Off on Week #49 – Stemming

Week #48 – Structured vs. unstructured data

Structured data is data that is in a form that can be used to develop statistical or machine learning models (typically a matrix where rows are records and columns are variables or features).

Comments Off on Week #48 – Structured vs. unstructured data

Week #47 – Feature engineering

In predictive modeling, a key step is to turn available data (which may come from varied sources and be messy) into an orderly matrix of rows (records to be predicted) and columns (predictor variables or features).

Comments Off on Week #47 – Feature engineering

Week #46 – Naive bayes classifier

A full Bayesian classifier is a supervised learning technique that assigns a class to a record by finding other records  with attributes just like it has, and finding the most prevalent class among them.

Comments Off on Week #46 – Naive bayes classifier

Week #45 – MapReduce

In computer science, MapReduce is a procedure that prepares data for parallel processing on multiple computers. 

Comments Off on Week #45 – MapReduce