Skip to main content
All projects
Machine learning project

Iris flower classifier

The iris dataset is a classic first dataset: 150 flowers from three species, each described by the length and width of its petals and sepals. You'll train a model on most of the flowers, test it on the ones it hasn't seen, and ask it about a flower of your own.

Skills you'll practise

  • Features and labels
  • Train/test split
  • k-nearest neighbours
  • Accuracy

Steps

  1. Step 1: Load the data

    Load the dataset and look at the feature names, the species names and the first few rows.

    Show a hint

    iris.data holds the measurements (the features), iris.target holds the species as numbers (the labels).

  2. Step 2: Hold some flowers back

    Split the data so the model trains on most flowers and is tested on ones it has never seen.

    Show a hint

    train_test_split(X, y, test_size=0.25, random_state=0, stratify=y) keeps a quarter back with all three species represented.

  3. Step 3: Train a model

    Create a KNeighborsClassifier and call .fit() on the training data.

  4. Step 4: Check how often it's right

    Predict the test flowers and compare the predictions with the true species using accuracy_score.

  5. Step 5: Ask about a new flower

    Make up four measurements and ask the model which species it thinks the flower is.

    Show a hint

    predict expects a list of flowers, so wrap a single flower in an extra pair of brackets: [[5.0, 3.4, 1.5, 0.2]].

Starter code

It already runs. The TODO comments mark where to start.

Open in playground
main.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

iris = load_iris()
X = iris.data
y = iris.target

print("Features:", iris.feature_names)
print("Species:", ", ".join(iris.target_names))
print("First flower:", X[0], "->", iris.target_names[y[0]])

# TODO: split the data, train a KNeighborsClassifier,
# and print its accuracy on the test flowers.

Example solution

One way to finish it. Have a go first; yours doesn't need to match.

Reveal the solution
main.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

iris = load_iris()
X = iris.data
y = iris.target

print("Features:", iris.feature_names)
print("Species:", ", ".join(iris.target_names))

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0, stratify=y
)

model = KNeighborsClassifier(n_neighbors=5)
model.fit(X_train, y_train)

predictions = model.predict(X_test)
print(f"Accuracy on unseen flowers: {accuracy_score(y_test, predictions):.0%}")

# sepal length, sepal width, petal length, petal width (cm)
new_flower = [[6.1, 2.9, 4.6, 1.4]]
guess = model.predict(new_flower)[0]
print("My flower looks like:", iris.target_names[guess])

Stretch goal

Try n_neighbors values from 1 to 15 in a loop and print the test accuracy for each. Does more neighbours always help?

Chapters that help