SLAI Natural Language Processing
Basic Information
- Instructor: Laura Biester, lbiester@umich.edu
- TA: Katrina Li, lik@carleton.edu
- Course Website
- Moodle Page
Course Description
Natural language processing (NLP) is the subfield of computer science that allows computers to represent, interpret, and generate text and spoken language. Even if you haven’t heard of NLP before, you have probably interacted with it – siri, alexa, your gmail spam detector and spell checkers are all examples of NLP deployed in the real world!
In this course, we will learn how to collect our own text datasets from sources like Twitter and Yelp. Then, we will explore how to clean our data and represent text in ways that computers can understand. Finally, we will explore a subset of the exciting applications of NLP, such as building a system that can identify the language of a sentence and using unstructured text to see how word usage has changed over time.
Schedule and Readings
There are assigned readings to complete before each day of the class. Unless otherwise stated, readings can be found on Moodle. The first readings for each day are required; assigned textbook chapters are optional and may be useful as a reference. On the first day, there is an optional reading from Foundations of Statistical Natural Language Processing.
| Date | Topic | Readings | Lab | |
|---|---|---|---|---|
| Class 1 | July 18 | Python and Tokens | [OPTIONAL] Foundations of Statistical Natural Language Processing Chapter 5: Collocations | Tokens and Collocations |
| Class 2 | July 19 | Text Generation with Markov Models | Math With Bad Drawings Chapter 20: The Book Shredders Language models: past, present, and future [OPTIONAL] Speech and Language Processing Chapter 3: N-gram Language Models |
Generating Movie Plots |
| Class 3 | July 20 | Language Identification with Naïve Bayes | Speech and Language Processing Chapter 4: Naive Bayes and Sentiment Classification | Language Identification |
| Class 4 | July 21 | Word Embeddings and Language Change | The Illustrated Word2vec Semantics derived automatically from language corpora contain human-like biases [OPTIONAL] Speech and Language Processing Chapter 6: Vector Semantics and Embeddings |
Measuring Differences in Word Usage with Embeddings |
Lab Assignments
During the NLP class at SLAI, we will work on lab assignments to practice programming and apply the concepts that you have learned about from the readings and the lectures. The labs are designed to be challenging, so don’t worry if you haven’t quite finished by the time the class is over, and be sure to ask questions!
Each lab will have a corresponding repl.it project for you to get started, which will contain data. For some labs, it will also include pre-defined utility functions in util.py or starter code in main.py.
If you have finished the main portion of the lab, you can choose an extension to implement under the extensions section at the end of the lab. Unless otherwise stated, these extensions can be completed in any order. These are meant to be less structured, and finishing all of them during class would be a great feat.
Pair Programming
You will work on the lab with a partner, using the pair programming methodology. Specifically, one student will act as the driver, while the other will be the navigator; a good explanation of those roles is available here. The instructors will announce to the class that it is time to switch drivers every 15 minutes.
Research Group
The students who take the NLP class during the first week of the program have the opportunity to complete a NLP research project that will span two weeks. Project ideas for the NLP research group is available here, and the project proposal guidelines are available here.