from dsc80_utils import *
Agenda 📆¶
- Text features.
- Bag of words.
- Cosine similarity.
- TF-IDF.
- Example: Presidential inaugural addresses 🎤.
Text features¶
Review: Regression and features¶
- In DSC 40A, our running example was to use regression to predict a data scientist's salary, given their GPA, years of experience, and years of education.
- After minimizing empirical risk to determine optimal parameters, $w_0^*, \dots, w_3^*$, we made predictions using:
$$\text{predicted salary} = w_0^* + w_1^* \cdot \text{GPA} + w_2^* \cdot \text{experience} + w_3^* \cdot \text{education}$$
- GPA, years of experience, and years of education are features – they represent a data scientist as a vector of numbers.
- e.g. Your feature vector may be [3.5, 1, 7].
- This approach requires features to be numeric.
Moving forward¶
Suppose we'd like to predict the sentiment of a piece of text from 1 to 10.
- 10: Very positive (happy).
- 1: Very negative (sad, angry).
Example:
Input: "DSC 80 is a pretty good class."
Output: 7.
We can frame this as a regression problem, but we can't directly use what we learned in 40A, because here our inputs are text, not numbers.
Text features¶
- Big question: How do we represent a text document as a feature vector of numbers?
- If we can do this, we can:
- use a text document as input in a regression or classification model (in a few lectures).
- quantify the similarity of two text documents (today).
Example: San Diego employee salaries¶
- Transparent California publishes the salaries of all City of San Diego employees.
- Let's look at the 2024 data (most recent available).
salaries = pd.read_csv('https://transcal.s3.amazonaws.com/public/export/san-diego-2024.csv')
salaries['Employee Name'] = salaries['Employee Name'].str.split().str[0] + ' Xxxx'
salaries.head()
| Employee Name | Job Title | Base Pay | Overtime Pay | ... | Year | Notes | Agency | Status | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | Carina Xxxx | Retirement Chief Investment Officer | 402834.40 | 0.0 | ... | 2024 | NaN | San Diego | FT |
| 1 | Gregg Xxxx | Retirement Administrator | 412903.58 | 0.0 | ... | 2024 | NaN | San Diego | FT |
| 2 | Eric Xxxx | Chief Operating Officer | 400708.40 | 0.0 | ... | 2024 | NaN | San Diego | FT |
| 3 | Mara Xxxx | City Attorney | 243583.77 | 0.0 | ... | 2024 | NaN | San Diego | FT |
| 4 | Todd Xxxx | Mayor | 243583.77 | 0.0 | ... | 2024 | NaN | San Diego | FT |
5 rows × 13 columns
Aside on privacy and ethics¶
- Even though the data we downloaded is publicly available, employee names still correspond to real people.
- Be careful when dealing with PII (personably identifiable information).
- Only work with the data that is needed for your analysis.
- Even when data is public, people have a reasonable right to privacy.
- Remember to think about the impacts of your work outside of your Jupyter Notebook.
Goal: Quantifying similarity¶
- Our goal is to describe, numerically, how similar two job titles are.
- For instance, our similarity metric should tell us that
'Deputy Fire Chief'and'Fire Battalion Chief'are more similar than'Deputy Fire Chief'and'City Attorney'.
- Idea: Two job titles are similar if they contain shared words, regardless of order. So, to measure the similarity between two job titles, we could count the number of words they share in common.
- Before we do this, we need to be confident that the job titles are clean and consistent – let's explore.
Exploring job titles¶
jobtitles = salaries['Job Title']
jobtitles.head()
0 Retirement Chief Investment Officer 1 Retirement Administrator 2 Chief Operating Officer 3 City Attorney 4 Mayor Name: Job Title, dtype: object
How many employees are in the dataset? How many unique job titles are there?
jobtitles.shape[0], jobtitles.nunique()
(14570, 660)
What are the most common job titles?
jobtitles.value_counts().iloc[:10]
Job Title
Police Officer Ii 1065
Management Intern 384
Student Intern 362
...
Grounds Maintenance Worker Ii 286
Police Detective 256
Fire Captain 244
Name: count, Length: 10, dtype: int64
Are there any missing job titles?
jobtitles.isna().sum()
0
Fortunately, no.
Canonicalization¶
Remember, our goal is ultimately to count the number of shared words between job titles. But before we start counting the number of shared words, we need to consider the following:
- Some job titles may have punctuation, like
'-'and'&', which may count as words when they shouldn't.'Assistant - Manager'and'Assistant Manager'should count as the same job title.
- Some job titles may have "glue" words, like
'to'and'the', which (we can argue) also shouldn't count as words.'Assistant To The Manager'and'Assistant Manager'should count as the same job title.
- If we just want to focus on the titles themselves, then perhaps roman numerals should be removed: that is,
'Police Officer Ii'and'Police Officer I'should count as the same job title.
Let's address the above issues. The process of converting job titles so that they are always represented the same way is called canonicalization.
Punctuation¶
Are there job titles with unnecessary punctuation that we can remove?
To find out, we can write a regular expression that looks for characters other than letters, numbers, and spaces.
We can use regular expressions with the
.strmethods we learned earlier in the quarter just by usingregex=True.
# Uses character class negation.
jobtitles.str.contains(r'[^A-Za-z0-9 ]', regex=True).sum()
1148
jobtitles[jobtitles.str.contains(r'[^A-Za-z0-9 ]', regex=True)].head()
96 Park & Recreation Director 536 Associate Engineer - Mechanical 565 Associate Engineer - Traffic 670 Associate Engineer - Electrical 744 Associate Engineer - Civil Name: Job Title, dtype: object
It seems like we should replace these pieces of punctuation with a single space.
"Glue" words¶
Are there job titles with "glue" words in the middle, such as 'Assistant To The Manager'?
To figure out if any titles contain the word 'to', we can't just do the following, because it will evaluate to True for job titles that have 'to' anywhere in them, even if not as a standalone word.
# Why are we converting to lowercase?
jobtitles.str.lower().str.contains('to').sum()
1806
Instead, we need to look for 'to' separated by word boundaries.
jobtitles.str.lower().str.contains(r'\bto\b', regex=True).sum()
13
jobtitles[jobtitles.str.lower().str.contains(r'\bto\b', regex=True)]
1585 Assistant To The Water Department Director
2116 Assistant To The Director
2449 Principal Assistant To City Attorney
...
6904 Assistant To The Director
8796 Assistant To The Chief Operating Officer
11101 Confidential Secretary To Mayor
Name: Job Title, Length: 13, dtype: object
We can look for other filler words too, like 'the' and 'for'.
jobtitles[jobtitles.str.lower().str.contains(r'\bthe\b', regex=True)]
1585 Assistant To The Water Department Director 2116 Assistant To The Director 4763 Assistant To The Director 6719 Assistant To The Director 6904 Assistant To The Director 8796 Assistant To The Chief Operating Officer Name: Job Title, dtype: object
jobtitles[jobtitles.str.lower().str.contains(r'\bfor\b', regex=True)]
228 Assistant For Community Outreach 4507 Assistant For Community Outreach Name: Job Title, dtype: object
We should probably remove these "glue" words.
Roman numerals (e.g. "Ii")¶
Lastly, let's try to identify job titles that have roman numerals at the end, like 'i' (1), 'ii' (2), 'iii' (3), or 'iv' (4). As before, we'll convert to lowercase first.
jobtitles[jobtitles.str.lower().str.contains(r'\bi+v?\b', regex=True)]
14 Police Officer Ii
20 Police Officer Ii
34 Police Officer Ii
...
14565 Water Systems Technician Iv
14568 Police Officer Ii
14569 Clerical Assistant Ii
Name: Job Title, Length: 6537, dtype: object
Let's get rid of those numbers, too.
Fixing punctuation and removing "glue" words and roman numerals¶
Let's put the preceeding three steps together and canonicalize job titles by:
- converting to lowercase,
- removing each occurrence of
'to','the', and'for', - replacing each character that is not a letter, digit, or space with a single space,
- replacing each sequence of roman numerals – either
'i','ii','iii', or'iv'at the end with nothing, and - replacing each sequence of multiple spaces with a single space.
jobtitles = (
jobtitles
.str.lower()
.str.replace(r'\bto\b|\bthe\b\|bfor\b', '', regex=True)
.str.replace(r'[^A-Za-z0-9 ]', ' ', regex=True)
.str.replace(r'\bi+v?\b', '', regex=True)
.str.replace(r' +', ' ', regex=True) # ' +' matches 1 or more occurrences of a space.
.str.strip() # Removes leading/trailing spaces if present.
)
jobtitles.sample(5)
14308 pool guard 5990 plumbing supervisor 135 deputy city attorney 11094 library assistant 11610 power plant operator Name: Job Title, dtype: object
(jobtitles == 'police officer').sum()
1268
Bag of words 💰¶
Text similarity¶
Recall, our idea is to measure the similarity of two job titles by counting the number of shared words between the job titles. How do we actually do that, for all of the job titles we have?
A counts matrix¶
Let's create a "counts" matrix, such that:
- there is 1 row per job title,
- there is 1 column per unique word that is used in job titles, and
- the value in row
titleand columnwordis the number of occurrences ofwordintitle.
Such a matrix might look like:
| senior | lecturer | teaching | professor | assistant | associate | |
|---|---|---|---|---|---|---|
| senior lecturer | 1 | 1 | 0 | 0 | 0 | 0 |
| assistant teaching professor | 0 | 0 | 1 | 1 | 1 | 0 |
| associate professor | 0 | 0 | 0 | 1 | 0 | 1 |
| senior assistant to the assistant professor | 1 | 0 | 0 | 1 | 2 | 0 |
Then, we can make statements like: "assistant teaching professor" is more similar to "associate professor" than to "senior lecturer".
Creating a counts matrix¶
First, we need to determine all words that are used across all job titles.
jobtitles.str.split()
0 [retirement, chief, investment, officer]
1 [retirement, administrator]
2 [chief, operating, officer]
...
14567 [assistant, fleet, technician]
14568 [police, officer]
14569 [clerical, assistant]
Name: Job Title, Length: 14570, dtype: object
# The .explode method concatenates the lists together into a single Series.
all_words = jobtitles.str.split().explode()
all_words
0 retirement
0 chief
0 investment
...
14568 officer
14569 clerical
14569 assistant
Name: Job Title, Length: 34583, dtype: object
Next, to determine the columns of our matrix, we need to find a list of all unique words used in titles. We can do this with np.unique, but value_counts shows us the distribution, which is interesting.
unique_words = all_words.value_counts()
unique_words
Job Title
police 2174
officer 1570
assistant 1345
...
ltd 1
termed 1
participant 1
Name: count, Length: 346, dtype: int64
Note that in unique_words.index, job titles are sorted by number of occurrences!
For each of the unique words that are used in job titles, we can count the number of occurrences of the word in each job title.
'deputy fire chief'contains the word'deputy'once, the word'fire'once, and the word'chief'once.'assistant managers assistant'contains the word'assistant'twice and the word'managers'once.
# Created using a dictionary to avoid a "DataFrame is highly fragmented" warning.
counts_dict = {}
for word in unique_words.index:
re_pat = fr'\b{word}\b'
counts_dict[word] = jobtitles.str.count(re_pat)
counts_df = pd.DataFrame(counts_dict).set_index(jobtitles)
counts_df.head()
| police | officer | assistant | engineer | ... | warehouse | ltd | termed | participant | |
|---|---|---|---|---|---|---|---|---|---|
| Job Title | |||||||||
| retirement chief investment officer | 0 | 1 | 0 | 0 | ... | 0 | 0 | 0 | 0 |
| retirement administrator | 0 | 0 | 0 | 0 | ... | 0 | 0 | 0 | 0 |
| chief operating officer | 0 | 1 | 0 | 0 | ... | 0 | 0 | 0 | 0 |
| city attorney | 0 | 0 | 0 | 0 | ... | 0 | 0 | 0 | 0 |
| mayor | 0 | 0 | 0 | 0 | ... | 0 | 0 | 0 | 0 |
5 rows × 346 columns
counts_df.shape
(14570, 346)
counts_df has one row for each employee, and one column for each unique word that is used in a job title.
Bag of words¶
- The bag of words model represents texts (e.g. job titles, sentences, documents) as vectors of word counts.
- The "counts" matrices we have worked with so far were created using the bag of words model.
- The bag of words model defines a vector space in $\mathbb{R}^{\text{number of unique words}}$.
- In the matrix on the previous slide, each row was a vector corresponding to a specific job title.
- It is called "bag of words" because it doesn't consider order.

Cosine similarity¶
Question: Which job titles are most similar to 'deputy fire chief'?¶
- Remember, our idea was to count the number of shared words between two job titles.
- We now have access to
counts_df, which contains a row vector for each job title.
- How can we use it to count the number of shared words between two job titles?
Counting shared words¶
To start, let's compare the row vectors for 'deputy fire chief' and 'fire battalion chief'.
dfc = counts_df.loc['deputy fire chief'].iloc[0]
dfc
police 0
officer 0
assistant 0
..
ltd 0
termed 0
participant 0
Name: deputy fire chief, Length: 346, dtype: int64
fbc = counts_df.loc['fire battalion chief'].iloc[0]
fbc
police 0
officer 0
assistant 0
..
ltd 0
termed 0
participant 0
Name: fire battalion chief, Length: 346, dtype: int64
We can stack these two vectors horizontally.
pair_counts = (
pd.concat([dfc, fbc], axis=1)
.sort_values(by=['deputy fire chief', 'fire battalion chief'], ascending=False)
.head(10)
.T
)
pair_counts
| fire | chief | deputy | battalion | ... | assistant | engineer | intern | civil | |
|---|---|---|---|---|---|---|---|---|---|
| deputy fire chief | 1 | 1 | 1 | 0 | ... | 0 | 0 | 0 | 0 |
| fire battalion chief | 1 | 1 | 0 | 1 | ... | 0 | 0 | 0 | 0 |
2 rows × 10 columns
'deputy fire chief' and 'fire battalion chief' have 2 shared words in common. One way to arrive at this result mathematically is by taking their dot product:
np.dot(pair_counts.iloc[0], pair_counts.iloc[1])
2
Recall: The dot product¶
- Recall, if $\vec{a} = \begin{bmatrix} a_1 & a_2 & ... & a_n \end{bmatrix}^T$ and $\vec{b} = \begin{bmatrix} b_1 & b_2 & ... & b_n \end{bmatrix}^T$ are two vectors, then their dot product $\vec{a} \cdot \vec{b}$ is defined as:
$$\vec{a} \cdot \vec{b} = a_1b_1 + a_2b_2 + ... + a_nb_n$$
- The dot product also has a geometric interpretation. If $|\vec{a}|$ and $|\vec{b}|$ are the $L_2$ norms (lengths) of $\vec{a}$ and $\vec{b}$, and $\theta$ is the angle between $\vec{a}$ and $\vec{b}$, then:
$$\vec{a} \cdot \vec{b} = |\vec{a}| |\vec{b}| \cos \theta$$
(source)- $\cos \theta$ is equal to its maximum value (1) when $\theta = 0$, i.e. when $\vec{a}$ and $\vec{b}$ point in the same direction.
Cosine similarity and bag of words¶
To measure the similarity between two word vectors, instead of just counting the number of shared words, we should compute their normalized dot product, also known as their cosine similarity.
$$\cos \theta = \boxed{\frac{\vec{a} \cdot \vec{b}}{|\vec{a}| | \vec{b}|}}$$
- If all elements in $\vec{a}$ and $\vec{b}$ are non-negative, then $\cos \theta$ ranges from 0 to 1.
- 🚨 Key idea: The larger $\cos \theta$ is, the more similar the two vectors are!
- It is important to normalize by the lengths of the vectors, otherwise texts with more words will have artificially high similarities with other texts.
Normalizing¶
$$\cos \theta = \boxed{\frac{\vec{a} \cdot \vec{b}}{|\vec{a}| | \vec{b}|}}$$
- Why can't we just use the dot product – that is, why must we divide by $|\vec{a}| | \vec{b}|$?
- Consider the following example:
| big | data | science | |
|---|---|---|---|
| big big big big data | 4 | 1 | 0 |
| big data science | 1 | 1 | 1 |
| science big data | 1 | 1 | 1 |
| Pair | Dot Product | Cosine Similarity |
|---|---|---|
| big data science and big big big big data | 5 | 0.7001 |
| big data science and science big data | 3 | 1 |
'big big big big data'has a large dot product with'big data science'just because it has the word'big'four times. But intuitively,'big data science'and'science big data'should be as similar as possible, since they're permutations of the same phrase.
- So, make sure to compute the cosine similarity – don't just use the dot product!
Note: Sometimes, you will see the cosine distance being used. It is the complement of cosine similarity:
$$\text{dist}(\vec{a}, \vec{b}) = 1 - \cos \theta$$
If $\text{dist}(\vec{a}, \vec{b})$ is small, the two word vectors are similar.
A recipe for computing similarities¶
Given a set of documents, to find the most similar text to one document $d$ in particular:
- Use the bag of words model to create a counts matrix, in which:
- there is 1 row per document,
- there is 1 column per unique word that is used across documents, and
- the value in row
docand columnwordis the number of occurrences ofwordindoc.
- Compute the cosine similarity between $d$'s row vector and all other documents' row vectors.
- The other document with the greatest cosine similarity is the most similar, under the bag of words model.
Example: Global warming 🌎¶
Consider the following three documents.
sentences = pd.Series([
'I really really want global peace',
'I must enjoy global warming',
'We must solve climate change'
])
sentences
0 I really really want global peace 1 I must enjoy global warming 2 We must solve climate change dtype: object
Let's represent each document using the bag of words model.
unique_words = sentences.str.split().explode().value_counts()
unique_words
I 2
really 2
global 2
..
solve 1
climate 1
change 1
Name: count, Length: 12, dtype: int64
counts_dict = {}
for word in unique_words.index:
re_pat = fr'\b{word}\b'
counts_dict[word] = sentences.str.count(re_pat)
counts_df = pd.DataFrame(counts_dict).set_index(sentences)
counts_df
| I | really | global | must | ... | We | solve | climate | change | |
|---|---|---|---|---|---|---|---|---|---|
| I really really want global peace | 1 | 2 | 1 | 0 | ... | 0 | 0 | 0 | 0 |
| I must enjoy global warming | 1 | 0 | 1 | 1 | ... | 0 | 0 | 0 | 0 |
| We must solve climate change | 0 | 0 | 0 | 1 | ... | 1 | 1 | 1 | 1 |
3 rows × 12 columns
Let's now find the cosine similarity between each pair of documents.
counts_df
| I | really | global | must | ... | We | solve | climate | change | |
|---|---|---|---|---|---|---|---|---|---|
| I really really want global peace | 1 | 2 | 1 | 0 | ... | 0 | 0 | 0 | 0 |
| I must enjoy global warming | 1 | 0 | 1 | 1 | ... | 0 | 0 | 0 | 0 |
| We must solve climate change | 0 | 0 | 0 | 1 | ... | 1 | 1 | 1 | 1 |
3 rows × 12 columns
def sim_pair(s1, s2):
return np.dot(s1, s2) / (np.linalg.norm(s1) * np.linalg.norm(s2))
# Look at the documentation of the .corr method to see how this works!
counts_df.T.corr(sim_pair)
| I really really want global peace | I must enjoy global warming | We must solve climate change | |
|---|---|---|---|
| I really really want global peace | 1.00 | 0.32 | 0.0 |
| I must enjoy global warming | 0.32 | 1.00 | 0.2 |
| We must solve climate change | 0.00 | 0.20 | 1.0 |
Issue: Bag of words only encodes the words that each document uses, not their meanings.
- "I really really want global peace" and "We must solve climate change" have similar meanings, but have no shared words, and thus a low cosine similarity.
- "I really really want global peace" and "I must enjoy global warming" have very different meanings, but a relatively high cosine similarity.
Pitfalls of the bag of words model¶
Remember, the key assumption underlying the bag of words model is that two documents are similar if they share many words in common.
- The bag of words model doesn't consider order.
- The job titles
'deputy fire chief'and'chief fire deputy'are treated as the same.
- The job titles
- The bag of words model doesn't consider the meaning of words.
'I love data science'and'I hate data science'share 75% of their words, but have very different meanings.
- The bag of words model treats all words as being equally important.
'deputy'and'fire'have the same importance, even though'fire'is probably more important in describing someone's job title.- Let's address this point.
TF-IDF¶
The importance of words¶
Issue: The bag of words model doesn't know which words are "important" in a document. Consider the following document:
How do we determine which words are important?
- Repetition of words indicates importance, but
- the most common words ("the", "of", "at") often don't have much meaning!
Goal: Find a way of quantifying the importance of a word in a document by balancing the above two factors, i.e. find the word that best summarizes a document.
Term frequency¶
- The term frequency of a word (term) $t$ in a document $d$, denoted $\text{tf}(t, d)$ is the proportion of words in document $d$ that are equal to $t$.
$$ \text{tf}(t, d)= \frac{\text{\# of occurrences of $t$ in $d$}}{\text{total \# of words in $d$}} $$
- Example: What is the term frequency of "basket" in the following document?
Answer: $\frac{2}{10} = \frac15$.
Intuition: Words that occur often within a document are important to the document's meaning.
- If $\text{tf}(t, d)$ is large, then word $t$ occurs often in $d$.
- If $\text{tf}(t, d)$ is small, then word $t$ does not occur often $d$.
- Issue: "the" also has a TF of $\frac15$, but it seems less important than "basket".
Inverse document frequency¶
- The inverse document frequency of a word $t$ in a set of documents $d_1, d_2, ...$ is
$$\text{idf}(t) = \log \left(\frac{\text{total \# of documents}}{\text{\# of documents in which $t$ appears}} \right)$$
- Example: What is the inverse document frequency of "basket" in the following three documents?
- "at the last game, the team scored basket after basket"
- "our team is on fire"
- "they only missed one basket"
- Answer: $\log \left(\frac{3}{2}\right) \approx 0.4055$.
- Intuition: If a word appears in every document (like "the", "of", "at"), it is probably not a good summary of any one document.
- If $\text{idf}(t)$ is large, then $t$ is rarely found in documents.
- If $\text{idf}(t)$ is small, then $t$ is commonly found in documents.
- Think of $\text{idf}(t)$ as the "rarity factor" of $t$ across documents – the larger $\text{idf}(t)$ is, the more rare $t$ is.
Intuition¶
$$\text{tf}(t, d) = \frac{\text{\# of occurrences of $t$ in $d$}}{\text{total \# of words in $d$}}$$
$$\text{idf}(t) = \log \left(\frac{\text{total \# of documents}}{\text{\# of documents in which $t$ appears}} \right)$$
Goal: Quantify how well word $t$ summarizes document $d$.
- If $\text{tf}(t, d)$ is small, then $t$ doesn't occur very often in $d$, so $t$ can't be a good summary of $d$.
- If $\text{idf}(t)$ is small, then $t$ occurs often amongst all documents, and so it is not a good summary of any one document.
- If $\text{tf}(t, d)$ and $\text{idf}(t)$ are both large, then $t$ occurs often in $d$ but rarely overall. This makes $t$ a good summary of document $d$.
Term frequency-inverse document frequency¶
The term frequency-inverse document frequency (TF-IDF) of word $t$ in document $d$ is the product:
$$ \begin{align*} \text{tfidf}(t, d) &= \text{tf}(t, d) \cdot \text{idf}(t) \\\ &= \frac{\text{\# of occurrences of $t$ in $d$}}{\text{total \# of words in $d$}} \cdot \log \left(\frac{\text{total \# of documents}}{\text{\# of documents in which $t$ appears}} \right) \end{align*} $$
- If $\text{tfidf}(t, d)$ is large, then $t$ is a good summary of $d$, because $t$ occurs often in $d$ but rarely across all documents.
- TF-IDF is a heuristic – it has no probabilistic justification.
- To know if $\text{tfidf}(t, d)$ is large for one particular word $t$, we need to compare it to $\text{tfidf}(t_i, d)$, for several different words $t_i$.
Computing TF-IDF¶
Question: What is the TF-IDF of "global" in the second sentence?
sentences
0 I really really want global peace 1 I must enjoy global warming 2 We must solve climate change dtype: object
Answer:
tf = sentences.iloc[1].count('global') / len(sentences.iloc[1].split())
tf
0.2
idf = np.log(len(sentences) / sentences.str.contains('global').sum())
idf
0.4054651081081644
tf * idf
0.08109302162163289
Question: Is this big or small? Is "global" the best summary of the second sentence?
TF-IDF of all words in all documents¶
On its own, the TF-IDF of a word in a document doesn't really tell us anything; we must compare it to TF-IDFs of other words in that same document.
sentences
0 I really really want global peace 1 I must enjoy global warming 2 We must solve climate change dtype: object
unique_words = np.unique(sentences.str.split().explode())
unique_words
array(['I', 'We', 'change', 'climate', 'enjoy', 'global', 'must', 'peace',
'really', 'solve', 'want', 'warming'], dtype=object)
tfidf_dict = {}
for word in unique_words:
re_pat = fr'\b{word}\b'
tf = sentences.str.count(re_pat) / sentences.str.split().str.len()
idf = np.log(len(sentences) / sentences.str.contains(re_pat).sum())
tfidf_dict[word] = tf * idf
tfidf = pd.DataFrame(tfidf_dict).set_index(sentences)
tfidf
| I | We | change | climate | ... | really | solve | want | warming | |
|---|---|---|---|---|---|---|---|---|---|
| I really really want global peace | 0.07 | 0.00 | 0.00 | 0.00 | ... | 0.37 | 0.00 | 0.18 | 0.00 |
| I must enjoy global warming | 0.08 | 0.00 | 0.00 | 0.00 | ... | 0.00 | 0.00 | 0.00 | 0.22 |
| We must solve climate change | 0.00 | 0.22 | 0.22 | 0.22 | ... | 0.00 | 0.22 | 0.00 | 0.00 |
3 rows × 12 columns
Interpreting TF-IDFs¶
display_df(tfidf, cols=12)
| I | We | change | climate | enjoy | global | must | peace | really | solve | want | warming | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| I really really want global peace | 0.07 | 0.00 | 0.00 | 0.00 | 0.00 | 0.07 | 0.00 | 0.18 | 0.37 | 0.00 | 0.18 | 0.00 |
| I must enjoy global warming | 0.08 | 0.00 | 0.00 | 0.00 | 0.22 | 0.08 | 0.08 | 0.00 | 0.00 | 0.00 | 0.00 | 0.22 |
| We must solve climate change | 0.00 | 0.22 | 0.22 | 0.22 | 0.00 | 0.00 | 0.08 | 0.00 | 0.00 | 0.22 | 0.00 | 0.00 |
The above DataFrame tells us that:
- the TF-IDF of
'really'in the first sentence is $\approx$ 0.37, - the TF-IDF of
'climate'in the second sentence is 0.
Note that there are two ways that $\text{tfidf}(t, d) = \text{tf}(t, d) \cdot \text{idf}(t)$ can be 0:
- If $t$ appears in every document, because then $\text{idf}(t) = \log (\frac{\text{\# documents}}{\text{\# documents}}) = \log(1) = 0$.
- If $t$ does not appear in document $d$, because then $\text{tf}(t, d) = \frac{0}{\text{len}(d)} = 0$.
The word that best summarizes a document is the word with the highest TF-IDF for that document:
display_df(tfidf, cols=12)
| I | We | change | climate | enjoy | global | must | peace | really | solve | want | warming | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| I really really want global peace | 0.07 | 0.00 | 0.00 | 0.00 | 0.00 | 0.07 | 0.00 | 0.18 | 0.37 | 0.00 | 0.18 | 0.00 |
| I must enjoy global warming | 0.08 | 0.00 | 0.00 | 0.00 | 0.22 | 0.08 | 0.08 | 0.00 | 0.00 | 0.00 | 0.00 | 0.22 |
| We must solve climate change | 0.00 | 0.22 | 0.22 | 0.22 | 0.00 | 0.00 | 0.08 | 0.00 | 0.00 | 0.22 | 0.00 | 0.00 |
tfidf.idxmax(axis=1)
I really really want global peace really I must enjoy global warming enjoy We must solve climate change We dtype: object
Look closely at the rows of tfidf – in documents 2 and 3, the max TF-IDF is not unique!
Example: Presidential inaugural addresses 🎤¶
Presidential inaugural addresses¶
Every four years, on January 20th, the incoming (or re-elected) President delivers an inaugural address just after taking the oath of office, in which they often share their vision for the country.
The data¶
from pathlib import Path
inaug_txt = Path('data') / 'inaugural_addresses.txt'
inaug = inaug_txt.read_text(encoding='utf-8')
len(inaug)
826494
The entire corpus (another word for "set of documents") is nearly 1 million characters long... let's not display it in our notebook.
print(inaug[:1500])
Presidential Inaugural Addresses, 1789 to 2025. Source: public-domain speech texts distributed via the NLTK 'inaugural' corpus (https://github.com/nltk/nltk_data, packages/corpora/inaugural.zip). Compiled for DSC 80 as a small, TF-IDF-friendly text corpus of real primary-source speeches. Each entry is separated by three asterisks. *** Inaugural Address Washington 1789 Fellow-Citizens of the Senate and of the House of Representatives: Among the vicissitudes incident to life no event could have filled me with greater anxieties than that of which the notification was transmitted by your order, and received on the 14th day of the present month. On the one hand, I was summoned by my Country, whose voice I can never hear but with veneration and love, from a retreat which I had chosen with the fondest predilection, and, in my flattering hopes, with an immutable decision, as the asylum of my declining years -- a retreat which was rendered every day more necessary as well as more dear to me by the addition of habit to inclination, and of frequent interruptions in my health to the gradual waste committed on it by time. On the other hand, the magnitude and difficulty of the trust to which the voice of my country called me, being sufficient to awaken in the wisest and most experienced of her citizens a distrustful scrutiny into his qualifications, could not but overwhelm with despondence one who (inheriting inferior endowments from nature and unpracticed in the duties of civil adminis
Each speech is separated by '***'.
speeches = inaug.split('\n***\n')[1:]
len(speeches)
60
Note that each "speech" currently contains other information, like the name of the president and the year of the address.
print(speeches[-1])
Inaugural Address Trump 2025 Thank you. Thank you very much, everybody. Wow. Thank you very, very much. Vice President Vance, Speaker Johnson, Senator Thune, Chief Justice Roberts, justices of the Supreme Court of the United States, President Clinton, President Bush, President Obama, President Biden, Vice President Harris, and my fellow citizens, the golden age of America begins right now. From this day forward, our country will flourish and be respected again all over the world. We will be the envy of every nation, and we will not allow ourselves to be taken advantage of any longer. During every single day of the Trump administration, I will, very simply, put America first. Our sovereignty will be reclaimed. Our safety will be restored. The scales of justice will be rebalanced. The vicious, violent, and unfair weaponization of the Justice Department and our government will end. And our top priority will be to create a nation that is proud, prosperous, and free. America will soon be greater, stronger, and far more exceptional than ever before. I return to the presidency confident and optimistic that we are at the start of a thrilling new era of national success. A tide of change is sweeping the country, sunlight is pouring over the entire world, and America has the chance to seize this opportunity like never before. But first, we must be honest about the challenges we face. While they are plentiful, they will be annihilated by this great momentum that the world is now witnessing in the United States of America. As we gather today, our government confronts a crisis of trust. For many years, a radical and corrupt establishment has extracted power and wealth from our citizens while the pillars of our society lay broken and seemingly in complete disrepair. We now have a government that cannot manage even a simple crisis at home while, at the same time, stumbling into a continuing catalogue of catastrophic events abroad. It fails to protect our magnificent, law-abiding American citizens but provides sanctuary and protection for dangerous criminals, many from prisons and mental institutions, that have illegally entered our country from all over the world. We have a government that has given unlimited funding to the defense of foreign borders but refuses to defend American borders or, more importantly, its own people. Our country can no longer deliver basic services in times of emergency, as recently shown by the wonderful people of North Carolina—who have been treated so badly—and other states who are still suffering from a hurricane that took place many months ago or, more recently, Los Angeles, where we are watching fires still tragically burn from weeks ago without even a token of defense. They’re raging through the houses and communities, even affecting some of the wealthiest and most powerful individuals in our country—some of whom are sitting here right now. They don’t have a home any longer. That’s interesting. But we can’t let this happen. Everyone is unable to do anything about it. That’s going to change. We have a public health system that does not deliver in times of disaster, yet more money is spent on it than any country anywhere in the world. And we have an education system that teaches our children to be ashamed of themselves—in many cases, to hate our country despite the love that we try so desperately to provide to them. All of this will change starting today, and it will change very quickly. My recent election is a mandate to completely and totally reverse a horrible betrayal and all of these many betrayals that have taken place and to give the people back their faith, their wealth, their democracy, and, indeed, their freedom. From this moment on, America’s decline is over. Our liberties and our nation’s glorious destiny will no longer be denied. And we will immediately restore the integrity, competency, and loyalty of America’s government. Over the past eight years, I have been tested and challenged more than any president in our 250-year history, and I’ve learned a lot along the way. The journey to reclaim our republic has not been an easy one—that, I can tell you. Those who wish to stop our cause have tried to take my freedom and, indeed, to take my life. Just a few months ago, in a beautiful Pennsylvania field, an assassin’s bullet ripped through my ear. But I felt then and believe even more so now that my life was saved for a reason. I was saved by God to make America great again. Thank you. Thank you. Thank you very much. That is why each day under our administration of American patriots, we will be working to meet every crisis with dignity and power and strength. We will move with purpose and speed to bring back hope, prosperity, safety, and peace for citizens of every race, religion, color, and creed. For American citizens, January 20th, 2025, is Liberation Day. It is my hope that our recent presidential election will be remembered as the greatest and most consequential election in the history of our country. As our victory showed, the entire nation is rapidly unifying behind our agenda with dramatic increases in support from virtually every element of our society: young and old, men and women, African Americans, Hispanic Americans, Asian Americans, urban, suburban, rural. And very importantly, we had a powerful win in all seven swing states and the popular vote, we won by millions of people. To the Black and Hispanic communities, I want to thank you for the tremendous outpouring of love and trust that you have shown me with your vote. We set records, and I will not forget it. I’ve heard your voices in the campaign, and I look forward to working with you in the years to come. Today is Martin Luther King Day. And his honor—this will be a great honor. But in his honor, we will strive together to make his dream a reality. We will make his dream come true. Thank you. Thank you. Thank you. National unity is now returning to America, and confidence and pride is soaring like never before. In everything we do, my administration will be inspired by a strong pursuit of excellence and unrelenting success. We will not forget our country, we will not forget our Constitution, and we will not forget our God. Can’t do that. Today, I will sign a series of historic executive orders. With these actions, we will begin the complete restoration of America and the revolution of common sense. It’s all about common sense. First, I will declare a national emergency at our southern border. All illegal entry will immediately be halted, and we will begin the process of returning millions and millions of criminal aliens back to the places from which they came. We will reinstate my Remain in Mexico policy. I will end the practice of catch and release. And I will send troops to the southern border to repel the disastrous invasion of our country. Under the orders I sign today, we will also be designating the cartels as foreign terrorist organizations. And by invoking the Alien Enemies Act of 1798, I will direct our government to use the full and immense power of federal and state law enforcement to eliminate the presence of all foreign gangs and criminal networks bringing devastating crime to US soil, including our cities and inner cities. As commander in chief, I have no higher responsibility than to defend our country from threats and invasions, and that is exactly what I am going to do. We will do it at a level that nobody has ever seen before. Next, I will direct all members of my cabinet to marshal the vast powers at their disposal to defeat what was record inflation and rapidly bring down costs and prices. The inflation crisis was caused by massive overspending and escalating energy prices, and that is why today I will also declare a national energy emergency. We will drill, baby, drill. America will be a manufacturing nation once again, and we have something that no other manufacturing nation will ever have—the largest amount of oil and gas of any country on earth—and we are going to use it. We’ll use it. We will bring prices down, fill our strategic reserves up again right to the top, and export American energy all over the world. We will be a rich nation again, and it is that liquid gold under our feet that will help to do it. With my actions today, we will end the Green New Deal, and we will revoke the electric vehicle mandate, saving our auto industry and keeping my sacred pledge to our great American autoworkers. In other words, you’ll be able to buy the car of your choice. We will build automobiles in America again at a rate that nobody could have dreamt possible just a few years ago. And thank you to the autoworkers of our nation for your inspiring vote of confidence. We did tremendously with their vote. I will immediately begin the overhaul of our trade system to protect American workers and families. Instead of taxing our citizens to enrich other countries, we will tariff and tax foreign countries to enrich our citizens. For this purpose, we are establishing the External Revenue Service to collect all tariffs, duties, and revenues. It will be massive amounts of money pouring into our Treasury, coming from foreign sources. The American dream will soon be back and thriving like never before. To restore competence and effectiveness to our federal government, my administration will establish the brand-new Department of Government Efficiency. After years and years of illegal and unconstitutional federal efforts to restrict free expression, I also will sign an executive order to immediately stop all government censorship and bring back free speech to America. Never again will the immense power of the state be weaponized to persecute political opponents—something I know something about. We will not allow that to happen. It will not happen again. Under my leadership, we will restore fair, equal, and impartial justice under the constitutional rule of law. And we are going to bring law and order back to our cities. This week, I will also end the government policy of trying to socially engineer race and gender into every aspect of public and private life. We will forge a society that is colorblind and merit-based. As of today, it will henceforth be the official policy of the United States government that there are only two genders: male and female. This week, I will reinstate any service members who were unjustly expelled from our military for objecting to the COVID vaccine mandate with full back pay. And I will sign an order to stop our warriors from being subjected to radical political theories and social experiments while on duty. It’s going to end immediately. Our armed forces will be freed to focus on their sole mission: defeating America’s enemies. Like in 2017, we will again build the strongest military the world has ever seen. We will measure our success not only by the battles we win but also by the wars that we end—and perhaps most importantly, the wars we never get into. My proudest legacy will be that of a peacemaker and unifier. That’s what I want to be: a peacemaker and a unifier. I’m pleased to say that as of yesterday, one day before I assumed office, the hostages in the Middle East are coming back home to their families. Thank you. America will reclaim its rightful place as the greatest, most powerful, most respected nation on earth, inspiring the awe and admiration of the entire world. A short time from now, we are going to be changing the name of the Gulf of Mexico to the Gulf of America, and we will restore the name of a great president, William McKinley, to Mount McKinley, where it should be and where it belongs. President McKinley made our country very rich through tariffs and through talent—he was a natural businessman—and gave Teddy Roosevelt the money for many of the great things he did, including the Panama Canal, which has foolishly been given to the country of Panama after the United Spates—the United States, I mean, think of this—spent more money than ever spent on a project before and lost 38,000 lives in the building of the Panama Canal. We have been treated very badly from this foolish gift that should have never been made, and Panama’s promise to us has been broken. The purpose of our deal and the spirit of our treaty has been totally violated. American ships are being severely overcharged and not treated fairly in any way, shape, or form. And that includes the United States Navy. And above all, China is operating the Panama Canal. And we didn’t give it to China. We gave it to Panama, and we’re taking it back. Above all, my message to Americans today is that it is time for us to once again act with courage, vigor, and the vitality of history’s greatest civilization. So, as we liberate our nation, we will lead it to new heights of victory and success. We will not be deterred. Together, we will end the chronic disease epidemic and keep our children safe, healthy, and disease-free. The United States will once again consider itself a growing nation—one that increases our wealth, expands our territory, builds our cities, raises our expectations, and carries our flag into new and beautiful horizons. And we will pursue our manifest destiny into the stars, launching American astronauts to plant the Stars and Stripes on the planet Mars. Ambition is the lifeblood of a great nation, and, right now, our nation is more ambitious than any other. There’s no nation like our nation. Americans are explorers, builders, innovators, entrepreneurs, and pioneers. The spirit of the frontier is written into our hearts. The call of the next great adventure resounds from within our souls. Our American ancestors turned a small group of colonies on the edge of a vast continent into a mighty republic of the most extraordinary citizens on Earth. No one comes close. Americans pushed thousands of miles through a rugged land of untamed wilderness. They crossed deserts, scaled mountains, braved untold dangers, won the Wild West, ended slavery, rescued millions from tyranny, lifted billions from poverty, harnessed electricity, split the atom, launched mankind into the heavens, and put the universe of human knowledge into the palm of the human hand. If we work together, there is nothing we cannot do and no dream we cannot achieve. Many people thought it was impossible for me to stage such a historic political comeback. But as you see today, here I am. The American people have spoken. I stand before you now as proof that you should never believe that something is impossible to do. In America, the impossible is what we do best. From New York to Los Angeles, from Philadelphia to Phoenix, from Chicago to Miami, from Houston to right here in Washington, DC, our country was forged and built by the generations of patriots who gave everything they had for our rights and for our freedom. They were farmers and soldiers, cowboys and factory workers, steelworkers and coal miners, police officers and pioneers who pushed onward, marched forward, and let no obstacle defeat their spirit or their pride. Together, they laid down the railroads, raised up the skyscrapers, built great highways, won two world wars, defeated fascism and communism, and triumphed over every single challenge that they faced. After all we have been through together, we stand on the verge of the four greatest years in American history. With your help, we will restore America promise and we will rebuild the nation that we love—and we love it so much. We are one people, one family, and one glorious nation under God. So, to every parent who dreams for their child and every child who dreams for their future, I am with you, I will fight for you, and I will win for you. We’re going to win like never before. Thank you. Thank you. Thank you. Thank you. In recent years, our nation has suffered greatly. But we are going to bring it back and make it great again, greater than ever before. We will be a nation like no other, full of compassion, courage, and exceptionalism. Our power will stop all wars and bring a new spirit of unity to a world that has been angry, violent, and totally unpredictable. America will be respected again and admired again, including by people of religion, faith, and goodwill. We will be prosperous, we will be proud, we will be strong, and we will win like never before. We will not be conquered, we will not be intimidated, we will not be broken, and we will not fail. From this day on, the United States of America will be a free, sovereign, and independent nation. We will stand bravely, we will live proudly, we will dream boldly, and nothing will stand in our way because we are Americans. The future is ours, and our golden age has just begun. Thank you. God bless America. Thank you all. Thank you. Thank you very much. Thank you very much. Thank you.
Let's extract just the speech text.
import re
def extract_struct(speech):
L = speech.strip().split('\n', maxsplit=3)
L[3] = L[3].replace('’', "'").replace('‘', "'") # standardize curly vs straight apostrophes
L[3] = re.sub(r"[^A-Za-z' ]", ' ', L[3]).lower()
return dict(zip(['speech', 'president', 'year', 'contents'], L))
speeches_df = pd.DataFrame(list(map(extract_struct, speeches)))
speeches_df
| speech | president | year | contents | |
|---|---|---|---|---|
| 0 | Inaugural Address | Washington | 1789 | fellow citizens of the senate and of the hous... |
| 1 | Inaugural Address | Washington | 1793 | fellow citizens i am again called upon by th... |
| 2 | Inaugural Address | Adams | 1797 | when it was first perceived in early times ... |
| ... | ... | ... | ... | ... |
| 57 | Inaugural Address | Trump | 2017 | chief justice roberts president carter pres... |
| 58 | Inaugural Address | Biden | 2021 | chief justice roberts vice president harris ... |
| 59 | Inaugural Address | Trump | 2025 | thank you thank you very much everybody wo... |
60 rows × 4 columns
Finding the most important words in each speech¶
Here, a "document" is a speech. We have 60 documents.
speeches_df
| speech | president | year | contents | |
|---|---|---|---|---|
| 0 | Inaugural Address | Washington | 1789 | fellow citizens of the senate and of the hous... |
| 1 | Inaugural Address | Washington | 1793 | fellow citizens i am again called upon by th... |
| 2 | Inaugural Address | Adams | 1797 | when it was first perceived in early times ... |
| ... | ... | ... | ... | ... |
| 57 | Inaugural Address | Trump | 2017 | chief justice roberts president carter pres... |
| 58 | Inaugural Address | Biden | 2021 | chief justice roberts vice president harris ... |
| 59 | Inaugural Address | Trump | 2025 | thank you thank you very much everybody wo... |
60 rows × 4 columns
A rough sketch of what we'll compute:
for each word t:
for each speech d:
compute tfidf(t, d)
unique_words = speeches_df['contents'].str.split().explode().value_counts().index
unique_words
Index(['the', 'of', 'and', 'to', 'in', 'a', 'our', 'we', 'that', 'be',
...
'businessman', 'tomb', 'crossed', 'scaled', 'braved', 'untold',
'unpredictable', 'admired', 'goodwill', 'houston'],
dtype='object', name='contents', length=9364)
💡 Pro-Tip: Using tqdm¶
This code takes a while to run, so we'll use the tqdm package to track its progress. (Install with mamba install tqdm if needed).
from tqdm.notebook import tqdm
tfidf_dict = {}
tf_denom = speeches_df['contents'].str.split().str.len()
# Wrap the sequence with `tqdm()` to display a progress bar
for word in tqdm(unique_words):
re_pat = fr' {word} ' # Imperfect pattern for speed.
tf = speeches_df['contents'].str.count(re_pat) / tf_denom
idf = np.log(len(speeches_df) / speeches_df['contents'].str.contains(re_pat).sum())
tfidf_dict[word] = tf * idf
0%| | 0/9364 [00:00<?, ?it/s]
tfidf = pd.DataFrame(tfidf_dict)
tfidf.sample(6, axis=0)
| the | of | and | to | ... | unpredictable | admired | goodwill | houston | |
|---|---|---|---|---|---|---|---|---|---|
| 12 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 40 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 4 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 56 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 49 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
6 rows × 9364 columns
Note that the TF-IDFs of many common words are all 0!
Summarizing speeches¶
By using idxmax, we can find the word with the highest TF-IDF in each speech.
summaries = tfidf.idxmax(axis=1)
summaries
0 immutable
1 arrive
2 pleasing
...
57 america
58 story
59 thank
Length: 60, dtype: object
What if we want to see the 10 words with the highest TF-IDFs, for each speech?
def ten_largest(row):
return ', '.join(row.index[row.argsort()][-10:])
keywords = tfidf.apply(ten_largest, axis=1)
keywords_df = pd.concat([
speeches_df['president'],
speeches_df['year'],
keywords
], axis=1)
display_df(keywords_df, rows=60)
| president | year | 0 | |
|---|---|---|---|
| 0 | Washington | 1789 | article, deliberations, rendered, pecuniary, your, peculiarly, qualifications, immutable, impressions, providential |
| 1 | Washington | 1793 | besides, injunctions, violated, knowingly, witnesses, previous, incurring, willingly, arrive, upbraidings |
| 2 | Adams | 1797 | benevolence, preference, esteem, legislature, amiable, habitual, virtuous, houses, legislatures, pleasing |
| 3 | Jefferson | 1801 | indeed, honest, error, him, trusted, principle, moments, intolerance, retire, thousandth |
| 4 | Jefferson | 1805 | press, expenses, comforts, licentiousness, defamation, enlighten, falsehood, covered, false, whatsoever |
| 5 | Madison | 1809 | rage, indulging, trespass, impression, assertions, fulfilling, invade, rendered, belligerent, improvements |
| 6 | Madison | 1813 | element, render, enemy, until, captives, prisoners, cruel, savage, massacre, british |
| 7 | Monroe | 1817 | dangers, connected, dependent, situated, duly, lakes, invasion, put, naval, trials |
| 8 | Monroe | 1821 | river, revenue, term, colonies, coast, fortifications, preceding, spain, occurrences, concluded |
| 9 | Adams | 1825 | general, internal, picture, contention, whatsoever, candid, performance, instituted, dissensions, union |
| 10 | Jackson | 1829 | mercifully, confederate, transcending, confound, customary, accountability, worth, defending, generally, diffidence |
| 11 | Jackson | 1833 | victorious, necessarily, impressed, exercise, intentions, preservation, union, constitutionally, inculcate, proportion |
| 12 | VanBuren | 1837 | opinions, unequaled, evidently, foreboding, designed, adherence, actual, occasionally, institutions, supposed |
| 13 | Harrison | 1841 | veto, used, constitution, remark, aristocracy, operations, conceive, executive, grant, roman |
| 14 | Polk | 1845 | extended, delegated, reunion, levying, majorities, compromises, revenue, her, union, texas |
| 15 | Taylor | 1849 | bestowal, surrounded, ambassadors, diligently, loftiest, loves, conflicting, affections, purity, vested |
| 16 | Pierce | 1853 | entirely, position, readily, counsels, apparent, regarded, preferment, consult, your, hardly |
| 17 | Buchanan | 1857 | convinced, possessions, slavery, speedily, squandering, kansas, question, agitation, territory, whilst |
| 18 | Lincoln | 1861 | fugitive, expressly, plainly, union, surrendered, lawfully, case, minority, secede, clause |
| 19 | Lincoln | 1865 | wringing, ascribe, address, altogether, answered, slaves, offense, wills, woe, offenses |
| 20 | Grant | 1869 | desirable, twenty, debt, payments, five, advisable, deal, specie, paying, dollar |
| 21 | Grant | 1873 | reformation, extension, him, steam, extermination, telegraph, proposition, domingo, santo, transit |
| 22 | Hayes | 1877 | emancipated, furtherance, parties, repeat, races, reform, nomination, complications, dispute, behalf |
| 23 | Garfield | 1881 | compulsory, smallest, jurisdiction, concerning, ballot, voters, suffrage, incumbents, notes, negro |
| 24 | Cleveland | 1885 | amity, claims, people's, application, appreciation, strife, commended, extravagance, partisan, yours |
| 25 | Harrison | 1889 | commercial, offices, list, cotton, electors, laws, friendly, european, methods, ballot |
| 26 | Cleveland | 1893 | paternalism, taxing, combinations, related, kindred, people's, governmental, activity, aggregations, frugality |
| 27 | McKinley | 1897 | merchant, revival, marine, convene, session, legislation, congress, revision, revenue, loans |
| 28 | McKinley | 1901 | advised, fast, intervention, executive, preparation, instructions, philippine, inhabitants, islands, cuba |
| 29 | Roosevelt | 1905 | unwasted, manlier, hardier, conditions, heritage, faced, tasks, problems, aright, regards |
| 30 | Taft | 1909 | amendment, bill, employees, canal, south, type, tariff, business, negro, interstate |
| 31 | Wilson | 1913 | comprehend, interpret, intimate, men's, safeguarding, sweep, inconceivable, studied, stirred, familiar |
| 32 | Wilson | 1917 | immediate, action, processes, politics, drawn, despite, currents, singular, counsel, wished |
| 33 | Harding | 1921 | wrought, proven, america, activities, unshaken, understanding, amid, normal, civilization, relationship |
| 34 | Coolidge | 1925 | financing, conditions, save, ought, unless, property, tax, stands, array, represents |
| 35 | Hoover | 1929 | criminals, reorganization, officials, mandates, controversies, amendment, liquor, ideals, eighteenth, enforcement |
| 36 | Roosevelt | 1933 | recovery, changers, languishes, critical, discipline, emergency, respects, stricken, leadership, helped |
| 37 | Roosevelt | 1937 | despair, millions, road, opportunism, epidemics, ruthless, aught, timidity, democracy, paint |
| 38 | Roosevelt | 1941 | midst, know, something, body, measured, three, stock, disruption, democracy, speaks |
| 39 | Roosevelt | 1945 | stout, faintness, schoolmaster, downward, gain, upward, mistakes, test, trend, learned |
| 40 | Truman | 1949 | democracy, agreement, technical, aided, communism, recovery, major, philosophy, peoples, program |
| 41 | Eisenhower | 1953 | mountains, defines, hold, spiritual, man's, korea, precepts, stamina, peoples, productivity |
| 42 | Eisenhower | 1957 | fate, freedom, honorably, seek, help, mr, divided, skills, peoples, strives |
| 43 | Kennedy | 1961 | finished, negotiate, deeds, pledge, dare, final, let, both, explore, sides |
| 44 | Johnson | 1965 | believers, trying, span, mars, harvest, rocket, shoulder, change, mastery, covenant |
| 45 | Nixon | 1969 | celebrate, wanting, worlds, caught, reaches, riders, third, brothers, rhetoric, voices |
| 46 | Nixon | 1973 | ashamed, america, shift, initiatives, let, gladly, abroad, america's, policies, role |
| 47 | Carter | 1977 | basic, together, mercy, bible, built, enhance, mistakes, thee, dream, micah |
| 48 | Reagan | 1981 | crosses, memorial, markers, row, monument, penalizes, americans, productivity, weapon, heroes |
| 49 | Reagan | 1985 | started, echoes, soviets, wouldn't, reduce, song, budget, senator, weapons, nuclear |
| 50 | Bush | 1989 | friends, mr, crucial, engagement, word, blowing, other's, door, don't, breeze |
| 51 | Clinton | 1993 | today, change, spring, capitol, sake, idea, americans, renewal, america, season |
| 52 | Clinton | 1997 | america, world's, journey, yes, streets, enough, st, promise, th, century |
| 53 | Bush | 2001 | everyone, angel, compassion, stakes, whirlwind, directs, rides, affirm, civility, story |
| 54 | Bush | 2005 | freedom's, fire, your, defended, america, americans, excuse, tyranny, freedom, america's |
| 55 | Obama | 2009 | journey, winter, earned, traveled, father, generation, storms, waters, jobs, icy |
| 56 | Obama | 2013 | thrives, requires, complete, until, evident, train, generation's, she, creed, journey |
| 57 | Trump | 2017 | transferring, politicians, stops, shine, everyone, jobs, dreams, we've, obama, america |
| 58 | Biden | 2021 | cry, we've, americans, democracy, america, uniting, virus, we're, let's, story |
| 59 | Trump | 2025 | immediately, win, america, sign, mckinley, you, back, panama, going, thank |
Aside: What if we remove the $\log$ from $\text{idf}(t)$?¶
Let's try it and see what happens.
tfidf_nl_dict = {}
tf_denom = speeches_df['contents'].str.split().str.len()
for word in tqdm(unique_words):
re_pat = fr' {word} ' # Imperfect pattern for speed.
tf = speeches_df['contents'].str.count(re_pat) / tf_denom
idf_nl = len(speeches_df) / speeches_df['contents'].str.contains(re_pat).sum()
tfidf_nl_dict[word] = tf * idf_nl
0%| | 0/9364 [00:00<?, ?it/s]
tfidf_nl = pd.DataFrame(tfidf_nl_dict)
tfidf_nl.head()
| the | of | and | to | ... | unpredictable | admired | goodwill | houston | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 0.08 | 0.05 | 0.03 | 0.03 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 1 | 0.10 | 0.08 | 0.01 | 0.04 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 2 | 0.07 | 0.06 | 0.06 | 0.03 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 3 | 0.08 | 0.06 | 0.05 | 0.04 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
| 4 | 0.07 | 0.05 | 0.04 | 0.04 | ... | 0.0 | 0.0 | 0.0 | 0.0 |
5 rows × 9364 columns
keywords_nl = tfidf_nl.apply(ten_largest, axis=1)
keywords_nl_df = pd.concat([
speeches_df['president'],
speeches_df['year'],
keywords_nl
], axis=1)
display_df(keywords_nl_df, rows=60)
| president | year | 0 | |
|---|---|---|---|
| 0 | Washington | 1789 | reflections, notification, preeminence, thence, renounce, aver, consulted, untried, of, the |
| 1 | Washington | 1793 | besides, injunctions, previous, violated, knowingly, witnesses, willingly, incurring, upbraidings, arrive |
| 2 | Adams | 1797 | sophistry, foreseen, to, legislatures, pleasing, amiable, habitual, and, of, the |
| 3 | Jefferson | 1801 | visionary, peaceable, unprovided, intentional, delights, to, and, of, thousandth, the |
| 4 | Jefferson | 1805 | to, whatsoever, and, of, licentiousness, defamation, enlighten, falsehood, covered, the |
| 5 | Madison | 1809 | inexpressibly, inadequacy, belligerent, exempted, manners, suppressing, regulates, to, of, the |
| 6 | Madison | 1813 | captives, extorted, disorganizing, savage, awaiting, insurrectional, hatchet, of, the, massacre |
| 7 | Monroe | 1817 | our, in, trials, duly, situated, lakes, and, to, of, the |
| 8 | Monroe | 1821 | croix, defrayed, spain, occurrences, concluded, in, and, to, of, the |
| 9 | Adams | 1825 | forest, aggravated, prepossessions, bounded, in, dissensions, to, and, of, the |
| 10 | Jackson | 1829 | confound, legible, induces, customary, cover, convinces, impenetrable, consist, of, the |
| 11 | Jackson | 1833 | victorious, imbibed, annihilation, lawgivers, deluge, encroaches, impair, link, of, the |
| 12 | VanBuren | 1837 | our, in, occasionally, unequaled, evidently, foreboding, to, and, of, the |
| 13 | Harrison | 1841 | monarchy, checked, observable, and, aristocracy, remark, to, roman, of, the |
| 14 | Polk | 1845 | consummate, liabilities, reunion, levying, to, and, compromises, of, texas, the |
| 15 | Taylor | 1849 | admonitions, bestowal, weightiest, onerous, surrounded, ambassadors, harmonize, to, of, the |
| 16 | Pierce | 1853 | unobtrusive, replete, a, hardly, to, consult, preferment, and, of, the |
| 17 | Buchanan | 1857 | inhabitant, residents, remainder, in, and, to, kansas, squandering, of, the |
| 18 | Lincoln | 1861 | suit, fugitives, frustrated, to, of, clause, fugitive, dissatisfied, secede, the |
| 19 | Lincoln | 1865 | prediction, wills, offense, widow, ventured, orphan, eighth, drop, offenses, woe |
| 20 | Grant | 1869 | stipulated, repudiator, farthing, retrenchment, feasibility, abeyance, contracting, moment's, the, dollar |
| 21 | Grant | 1873 | aborigines, proposition, of, the, extermination, telegraph, steam, transit, domingo, santo |
| 22 | Hayes | 1877 | piety, constraint, september, in, to, and, complications, nomination, of, the |
| 23 | Garfield | 1881 | ballot, voters, and, compulsory, negro, smallest, notes, incumbents, of, the |
| 24 | Cleveland | 1885 | workshop, busy, subserviency, supplanted, faculty, freedmen, harmoniously, and, of, the |
| 25 | Harrison | 1889 | residing, districts, mill, applicants, contentions, to, ballot, and, of, the |
| 26 | Cleveland | 1893 | kindred, cheapness, wily, misappropriation, to, and, of, aggregations, the, frugality |
| 27 | McKinley | 1897 | convene, merchant, to, voted, foremost, revision, and, loans, of, the |
| 28 | McKinley | 1901 | defenders, maketh, to, islands, and, of, philippine, instructions, the, cuba |
| 29 | Roosevelt | 1905 | hardier, wither, boastfulness, vainglory, acknowledgment, giver, bygone, the, regards, aright |
| 30 | Taft | 1909 | philippines, antitrust, to, and, lock, negro, type, of, the, interstate |
| 31 | Wilson | 1913 | sanitary, and, of, the, inconceivable, sweep, safeguarding, stirred, familiar, studied |
| 32 | Wilson | 1917 | humors, sternly, audience, provincials, quick, delegation, of, and, the, wished |
| 33 | Harding | 1921 | staggering, acclaim, crave, spiritually, unselfishness, hateful, materially, and, of, the |
| 34 | Coolidge | 1925 | experiences, deliverance, financing, polls, to, and, array, of, represents, the |
| 35 | Hoover | 1929 | mandates, officials, to, stimulating, abounding, and, liquor, of, the, eighteenth |
| 36 | Roosevelt | 1933 | uneconomical, planning, stubbornness, helped, of, changers, languishes, critical, the, stricken |
| 37 | Roosevelt | 1937 | stagnation, fashioning, paint, the, of, timidity, ruthless, epidemics, opportunism, aught |
| 38 | Roosevelt | 1941 | function, decisively, clarity, ebbing, surging, prophecy, downfall, of, speaks, the |
| 39 | Roosevelt | 1945 | untroubled, trend, ostriches, dogs, manger, smoothly, emerson, stout, downward, anguished |
| 40 | Truman | 1949 | to, of, and, partners, regime, philosophy, developments, major, the, technical |
| 41 | Eisenhower | 1953 | in, we, to, and, stamina, korea, productivity, precepts, of, the |
| 42 | Eisenhower | 1957 | tormented, russia, needy, billion, desperation, bartered, and, of, the, strives |
| 43 | Kennedy | 1961 | subversion, casting, invective, huts, explore, embattled, communists, overburdened, of, the |
| 44 | Johnson | 1965 | tower, surroundings, companions, and, the, covenant, harvest, shoulder, rocket, mastery |
| 45 | Nixon | 1969 | in, we, to, of, riders, worlds, reaches, caught, the, rhetoric |
| 46 | Nixon | 1973 | cook, peking, gladly, flimsy, moscow, to, of, the, initiatives, shift |
| 47 | Carter | 1977 | o, milestone, emulation, earns, ensured, politically, craving, affirmation, unchanging, micah |
| 48 | Reagan | 1981 | weapon, productivity, and, markers, penalizes, memorial, monument, row, crosses, the |
| 49 | Reagan | 1985 | of, and, wouldn't, kill, echoes, started, soviets, mathias, infirm, the |
| 50 | Bush | 1989 | a, crucial, engagement, and, vow, door, the, other's, blowing, breeze |
| 51 | Clinton | 1993 | broadcast, instantaneously, tobillions, mobile, magical, devastates, myriad, eroded, and, the |
| 52 | Clinton | 1997 | repairers, breach, tools, bernardin, explored, to, our, and, of, the |
| 53 | Bush | 2001 | insignificant, scapegoats, angel, civility, and, story, stakes, directs, rides, whirlwind |
| 54 | Bush | 2005 | wheels, runs, eventual, murder, background, viewpoint, unwanted, and, of, the |
| 55 | Obama | 2009 | whip, khe, normandy, we, our, to, of, and, icy, the |
| 56 | Obama | 2013 | militias, lucky, to, of, we, our, and, the, generation's, train |
| 57 | Trump | 2017 | redistributed, the, and, we've, obama, politicians, stops, shine, transferring, trillions |
| 58 | Biden | 2021 | versus, disagreement, extremism, doesn't, pandemic, shoes, we're, uniting, virus, let's |
| 59 | Trump | 2025 | los, angeles, i've, hispanic, nobody, drill, the, and, panama, mckinley |
The role of $\log$ in $\text{idf}(t)$¶
$$ \begin{align*} \text{tfidf}(t, d) &= \text{tf}(t, d) \cdot \text{idf}(t) \\\ &= \frac{\text{\# of occurrences of $t$ in $d$}}{\text{total \# of words in $d$}} \cdot \log \left(\frac{\text{total \# of documents}}{\text{\# of documents in which $t$ appears}} \right) \end{align*} $$
- Remember, for any positive input $x$, $\log(x)$ is (much) smaller than $x$.
- In $\text{idf}(t)$, the $\log$ "dampens" the impact of the ratio $\frac{\text{\# documents}}{\text{\# documents with $t$}}$.
- If a word is very common, the ratio will be close to 1. The log of the ratio will be close to 0.
(1000 / 999)
1.001001001001001
np.log(1000 / 999)
0.001000500333583622
- If a word is very common (e.g. 'the'), removing the log multiplies the statistic by a large factor.
- If a word is very rare, the ratio will be very large. However, for instance, a word being seen in 2 out of 50 documents is not very different than being seen in 2 out of 500 documents (it is very rare in both cases), and so $\text{idf}(t)$ should be similar in both cases.
(50 / 2)
25.0
(500 / 2)
250.0
np.log(50 / 2)
3.2188758248682006
np.log(500 / 2)
5.521460917862246
Question 🤔
From the Fall 23 final: Consider the following corpus:
Document number Content
1 yesterday rainy today sunny
2 yesterday sunny today sunny
3 today rainy yesterday today
4 yesterday yesterday today today
- Using a bag-of-words representation, which two documents have the largest dot product?
- Using a bag-of-words representation, what is the cosine similarity between documents 2 and 3?
- Which words have a TF-IDF score of 0 for all four documents?
Summary, next time¶
Summary¶
- One way to turn documents, like
'deputy fire chief', into feature vectors, is to count the number of occurrences of each word in the document, ignoring order. This is done using the bag of words model. - To measure the similarity of two documents under the bag of words model, compute the cosine similarity of their two word vectors.
- Term frequency-inverse document frequency (TF-IDF) is a statistic that tries to quantify how important a word (term) is to a document. It balances:
- how often a word appears in a particular document, $\text{tf}(t, d)$, with
- how often a word appears across documents, $\text{idf}(t)$.
- For a given document, the word with the highest TF-IDF is thought to "best summarize" that document.
Next time¶
Modeling and feature engineering.