Word-frequency counter
Writers, search engines and spam filters all count words. You'll clean up a paragraph, count every word, skip the tiny ones like "the" and "and", and draw a little bar chart of the favourites in plain text.
Skills you'll practise
- Strings
- Dictionaries
- Sorting
- Loops
Steps
Step 1: Split the text into words
Lower-case the text and split it on spaces so "Garden" and "garden" count as the same word.
Show a hintHide the hint
text.lower().split()splits on any run of spaces or new lines.Step 2: Strip the punctuation
Remove commas, full stops and quotes from the ends of each word.
Show a hintHide the hint
word.strip(string.punctuation)removes punctuation from both ends. Skip anything that ends up empty.Step 3: Count each word
Use a dictionary where each key is a word and each value is how many times you've seen it.
Show a hintHide the hint
counts[word] = counts.get(word, 0) + 1works whether or not the word is already there.Step 4: Skip the filler words
Make a set of common words like "the", "a" and "and", and leave them out of the count.
Step 5: Show the top ten
Sort the words by count, biggest first, and print each one with a bar of
#characters.Show a hintHide the hint
sorted(counts.items(), key=lambda pair: (-pair[1], pair[0]))sorts by count, then alphabetically to break ties.
Starter code
It already runs. The TODO comments mark where to start.
import string
text = """
Our school garden started as one bed of tomatoes. Then someone planted beans,
and the beans climbed the fence, and the fence became a wall of green.
Now the garden has herbs, flowers and a bench where people sit and read.
Every spring, new students ask what they can plant, and the garden grows again.
"""
# TODO: lower-case the text, split it into words, strip punctuation,
# and count how many times each word appears.
words = text.split()
print(len(words), "words")
Example solution
One way to finish it. Have a go first; yours doesn't need to match.
Reveal the solutionHide the solution
import string
text = """
Our school garden started as one bed of tomatoes. Then someone planted beans,
and the beans climbed the fence, and the fence became a wall of green.
Now the garden has herbs, flowers and a bench where people sit and read.
Every spring, new students ask what they can plant, and the garden grows again.
"""
STOP_WORDS = {"a", "an", "and", "as", "of", "the", "then", "they", "what", "where", "can", "has", "now", "our"}
def clean_words(text):
words = []
for raw in text.lower().split():
word = raw.strip(string.punctuation)
if word:
words.append(word)
return words
counts = {}
for word in clean_words(text):
if word in STOP_WORDS:
continue
counts[word] = counts.get(word, 0) + 1
top = sorted(counts.items(), key=lambda pair: (-pair[1], pair[0]))[:10]
for word, count in top:
print(f"{word:<10} {'#' * count} {count}")
Stretch goal
Paste in a longer piece of your own writing, or try collections.Counter and its most_common() method and compare it with your version.