Skip to main content
All projects
Python project

Word-frequency counter

Writers, search engines and spam filters all count words. You'll clean up a paragraph, count every word, skip the tiny ones like "the" and "and", and draw a little bar chart of the favourites in plain text.

Skills you'll practise

  • Strings
  • Dictionaries
  • Sorting
  • Loops

Steps

  1. Step 1: Split the text into words

    Lower-case the text and split it on spaces so "Garden" and "garden" count as the same word.

    Show a hint

    text.lower().split() splits on any run of spaces or new lines.

  2. Step 2: Strip the punctuation

    Remove commas, full stops and quotes from the ends of each word.

    Show a hint

    word.strip(string.punctuation) removes punctuation from both ends. Skip anything that ends up empty.

  3. Step 3: Count each word

    Use a dictionary where each key is a word and each value is how many times you've seen it.

    Show a hint

    counts[word] = counts.get(word, 0) + 1 works whether or not the word is already there.

  4. Step 4: Skip the filler words

    Make a set of common words like "the", "a" and "and", and leave them out of the count.

  5. Step 5: Show the top ten

    Sort the words by count, biggest first, and print each one with a bar of # characters.

    Show a hint

    sorted(counts.items(), key=lambda pair: (-pair[1], pair[0])) sorts by count, then alphabetically to break ties.

Starter code

It already runs. The TODO comments mark where to start.

Open in playground
main.py
import string

text = """
Our school garden started as one bed of tomatoes. Then someone planted beans,
and the beans climbed the fence, and the fence became a wall of green.
Now the garden has herbs, flowers and a bench where people sit and read.
Every spring, new students ask what they can plant, and the garden grows again.
"""

# TODO: lower-case the text, split it into words, strip punctuation,
# and count how many times each word appears.
words = text.split()
print(len(words), "words")

Example solution

One way to finish it. Have a go first; yours doesn't need to match.

Reveal the solution
main.py
import string

text = """
Our school garden started as one bed of tomatoes. Then someone planted beans,
and the beans climbed the fence, and the fence became a wall of green.
Now the garden has herbs, flowers and a bench where people sit and read.
Every spring, new students ask what they can plant, and the garden grows again.
"""

STOP_WORDS = {"a", "an", "and", "as", "of", "the", "then", "they", "what", "where", "can", "has", "now", "our"}


def clean_words(text):
    words = []
    for raw in text.lower().split():
        word = raw.strip(string.punctuation)
        if word:
            words.append(word)
    return words


counts = {}
for word in clean_words(text):
    if word in STOP_WORDS:
        continue
    counts[word] = counts.get(word, 0) + 1

top = sorted(counts.items(), key=lambda pair: (-pair[1], pair[0]))[:10]

for word, count in top:
    print(f"{word:<10} {'#' * count} {count}")

Stretch goal

Paste in a longer piece of your own writing, or try collections.Counter and its most_common() method and compare it with your version.

Chapters that help