Iris flower classifier
The iris dataset is a classic first dataset: 150 flowers from three species, each described by the length and width of its petals and sepals. You'll train a model on most of the flowers, test it on the ones it hasn't seen, and ask it about a flower of your own.
Skills you'll practise
- Features and labels
- Train/test split
- k-nearest neighbours
- Accuracy
Steps
Step 1: Load the data
Load the dataset and look at the feature names, the species names and the first few rows.
Show a hintHide the hint
iris.dataholds the measurements (the features),iris.targetholds the species as numbers (the labels).Step 2: Hold some flowers back
Split the data so the model trains on most flowers and is tested on ones it has never seen.
Show a hintHide the hint
train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)keeps a quarter back with all three species represented.Step 3: Train a model
Create a
KNeighborsClassifierand call.fit()on the training data.Step 4: Check how often it's right
Predict the test flowers and compare the predictions with the true species using
accuracy_score.Step 5: Ask about a new flower
Make up four measurements and ask the model which species it thinks the flower is.
Show a hintHide the hint
predictexpects a list of flowers, so wrap a single flower in an extra pair of brackets:[[5.0, 3.4, 1.5, 0.2]].
Starter code
It already runs. The TODO comments mark where to start.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
iris = load_iris()
X = iris.data
y = iris.target
print("Features:", iris.feature_names)
print("Species:", ", ".join(iris.target_names))
print("First flower:", X[0], "->", iris.target_names[y[0]])
# TODO: split the data, train a KNeighborsClassifier,
# and print its accuracy on the test flowers.
Example solution
One way to finish it. Have a go first; yours doesn't need to match.
Reveal the solutionHide the solution
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
iris = load_iris()
X = iris.data
y = iris.target
print("Features:", iris.feature_names)
print("Species:", ", ".join(iris.target_names))
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0, stratify=y
)
model = KNeighborsClassifier(n_neighbors=5)
model.fit(X_train, y_train)
predictions = model.predict(X_test)
print(f"Accuracy on unseen flowers: {accuracy_score(y_test, predictions):.0%}")
# sepal length, sepal width, petal length, petal width (cm)
new_flower = [[6.1, 2.9, 4.6, 1.4]]
guess = model.predict(new_flower)[0]
print("My flower looks like:", iris.target_names[guess])
Stretch goal
Try n_neighbors values from 1 to 15 in a loop and print the test accuracy for each. Does more neighbours always help?