Testing Models Like Human Subjects

We already tested human participants on psychophysics tasks. Now we test artificial neural networks on the same task images and ask: do the models behave like human subjects?

CVR Summer School 2026

ANN notebooks

Use these launch buttons to open the original notebooks in Google Colab. The tutorial content below is the original version. For each demo, make a copy of these original files to your Drive.

← Back to tutorial hub
AlexNet demoanns/AlexNet_demo.ipynb
Open in Colab
Facial emotionanns/Facial_emotion.ipynb
Open in Colab
Object 2AFCanns/obj_2afc.ipynb
Open in Colab
Drawing 2AFCanns/draw_2afc.ipynb
Open in Colab
Ratingsanns/ratings.ipynb
Open in Colab
N-backanns/n_back.ipynb
Open in Colab
Human vs ANN comparisoncompare_humans_vs_anns.ipynb
Open in Colab

1. Goal of this tutorial

This tutorial explains how to make the models perform the same task as humans. The main idea is very simple:

Humans did a task → Models do the same task → We compare them

In the psychophysics experiments, a human participant saw images and gave responses. In the ANN notebooks, a model will see the same images and perform the same task as the human participants, producing the similar kind of responses. We call this the "model output".

Big idea: we treat the model like any other human subject. If the model captures relevant functions used by the human brain, it is probably a good approximation of the behavioural measurement, and in turn produces similar answers to the humans (for example, finding some images easier than others to judge and making mistakes for similar images). This allows us to gain insights into the underlying computations used by the human visual system.

2. What does it mean to "treat a model as a human subject"?

A human subject sees an image and produces a response. A model also receives an image and produces an output. The code makes these two things comparable.

Image → Human → Human response
Image → Model → Model output

The model output depends on the task. In a recognition task, the model may output probabilities for object categories. In a rating task, the model outputs a score. In an n-back task, the model outputs a memorability score for each image.

Important: the model does not literally press keys or move a mouse. Instead, the notebook converts the model's prediction into a value that can be compared with the human data.

3. Demo notebooks

Before running model experiments with "task-matched" notebooks, we will learn what happens in these experiments using demo notebooks. These demo notebooks are simpler than the task-matched notebooks, and they show how to load a pretrained model, how give it one image, and how read the model prediction.

Important distinction: the demo notebooks test models on the kind of task they were originally trained for. The task-matched notebooks test models in the same format as the human experiments.

Why do we include demo notebooks?

The demo notebooks are useful because they introduce the basic model workflow:

Notebook What the model was trained to do What goes in What comes out
AlexNet_demo.ipynb AlexNet was trained for image classification. It learns to look at an image and predict what object category is present. One image. A list of predicted object labels, usually ordered from most likely to least likely.
Facial_emotion.ipynb The facial emotion model was trained to classify facial expressions. It learns to look at a face and predict an emotion label. One face image. A list of predicted emotion labels and confidence scores.

AlexNet demo

In the AlexNet demo, the notebook loads a pretrained image-classification model, prepares one image, runs the image through the model, and produces the 5 most likely object categories.

One image → AlexNet → Object label predictions

The output is not a human 2AFC response yet. It is simply the model saying which object categories it thinks are most likely.

model = models.alexnet(pretrained=True)
model.eval()

with torch.no_grad():
    output = model(input_image)

_, indices = torch.topk(output, 5)

In simple terms, this asks: What object does the pretrained model think is in this image?

Facial emotion demo

In the facial emotion demo, a pretrained face-expression model is used. The notebook gives the model one face image and produces the most likely emotion label.

One face image → Facial emotion model → Emotion labels + scores
clf = pipeline(
    "image-classification",
    model="trpakov/vit-face-expression"
)

results = clf(img)

In simple terms, this asks: What emotion does the pretrained model think this face shows?

After understanding this basic idea, the task-matched notebooks extend the same logic to the full psychophysics tasks.

Demo notebook output: demo notebooks do not save the main human-comparison CSV files. They mainly show predictions for one example image.

4. Task-matched notebook guide

Notebook Human task What the model does Output file
obj_2afc.ipynb Object 2AFC: humans choose the correct object from two options. A recognition model predicts object probabilities. The code converts those probabilities into 2AFC-style image accuracy. results/obj_2afc.csv
draw_2afc.ipynb Drawing 2AFC: humans choose the correct object from two options, but the image is a drawing or sketch. The same recognition model is tested on drawings. The code again converts probabilities into 2AFC-style image accuracy. results/draw_2afc.csv
ratings.ipynb Ratings: humans rate face images. A face-rating model predicts one score for each face image. results/ratings.csv
n_back.ipynb N-back: humans perform a memory-related task with images. A memorability model predicts one memorability score for each image. results/n_back.csv
All notebooks save a simple CSV table with one row per image. This makes it easy to compare model outputs with human behavior image by image.

5. Shared code: loading images in the right order

The file utils.py contains helper code used by the notebooks. This keeps the notebooks shorter and easier to read.

Sorting images

Images are named like this:

im1.png
im2.png
im3.png
...

For face ratings, the images are .jpg files:

im1.jpg
im2.jpg
im3.jpg
...

The helper function sort_images extracts the number from each file name and sorts the images in numerical order.

def sort_images(x):
    xs = []
    for im in x:
        t = int(os.path.splitext(os.path.basename(im))[0].replace("im", ""))
        xs.append(t)
    return list(sorted(xs))

This matters because we need the model result for im12.png to line up with the human result for im12.png.

Preparing images for recognition models

The recognition model needs images to be resized and normalized before it can use them. The helper code resizes images to 224 × 224 and applies the same normalization expected by common image models.

im_transforms = transforms.Compose([
    transforms.Resize((224, 224)),
    transforms.Lambda(lambda x: (x / 255.0)),
    transforms.Normalize(
        mean=[0.485, 0.456, 0.406],
        std=[0.229, 0.224, 0.225]
    ),
])

The helper function make_val_loader loads all task images, applies these transformations, stacks them into one tensor, and creates a DataLoader.

Simple version: the helper code makes sure images are loaded, sorted, resized, normalized, and ready for the model.

6. Object and drawing 2AFC notebooks

The notebooks obj_2afc.ipynb and draw_2afc.ipynb use the same logic. The difference is the image folder:

What is a 2AFC task?

2AFC means two-alternative forced choice. The participant chooses between two possible answers.

For example:

Image of a bear → Choice: bear or dog? → Correct answer: bear

The model does not choose by clicking a button. Instead, it gives a probability for each object class. The code uses these probabilities to ask which option the model would prefer.

Model used in both 2AFC notebooks

The notebooks currently load AlexNet. AlexNet starts as a pretrained image-classification model. Then the final layer is changed so the model predicts 10 task categories. Finally, fine-tuned weights are loaded from checkpoints/alexnet_fine_tuned.pth.

model = models.alexnet(pretrained=True)
weights = torch.load("checkpoints/alexnet_fine_tuned.pth")

model.classifier[6] = nn.Linear(model.classifier[6].in_features, 10)
model.load_state_dict(weights)
model.eval()

The notebooks also include commented code for other possible models, such as ResNet50 and ConvNeXt.

Step 1: Load task images and labels

For object 2AFC:

data_loader = make_val_loader("../images/obj_2AFC", 8)

labels = [0, 0, 1, 1, 2, 2, 2, 3, 4, 4, 5, 6, 6, 7, 7, 8]

For drawing 2AFC:

data_loader = make_val_loader("../images/draw_2AFC", 4)

labels = [0, 1, 4, 5]

The labels tell the code which object category is correct for each image. The labels used in these tasks are defined in the metadata README: README_metadata_tests.md .

Step 2: Get model probabilities

The model is run on every image. Its output is in the form of probabilities per object category class.

probabilities = []

with torch.no_grad():
    for inputs in data_loader:
        inputs = inputs[0]
        output = model(inputs)
        output = torch.nn.functional.softmax(output, dim=-1)
        probabilities.append(output.detach().cpu().numpy())

probabilities = np.concatenate(probabilities, axis=0)

After this step, each image has 10 probabilities: one probability for each possible object category.

Step 3: Check normal classification accuracy

The notebook first checks the usual model accuracy. This asks whether the model's highest-probability category is the correct category.

preds = np.argmax(probabilities, axis=1)
(preds == labels).mean()

This is useful, but it is not exactly the human 2AFC task. Humans choose between two options, not all 10 categories at once.

Step 4: Convert probabilities into 2AFC-like accuracy

This is the key step. The code compares the model probability for the correct answer with the probability for each possible wrong answer.

For each image, the model checks how likely it would have chosen the correct category label compared to picking one of the wrong category labels:

score =
    probability_correct_category /
    (probability_correct_category + probability_wrong_category)

This gives a model 2AFC score. If the model gives more probability to the correct answer than to the wrong answer, the score is high. If it gives more probability to the wrong answer, the score is low. The score is computed across all possible distractors and averaged to produce a single score per image (called image_scores below).

image_scores, model_choices = create_image_level_accuracies(
    probabilities,
    labels=labels
)

The output image_scores gives one 2AFC-style accuracy score per image. The output model_choices is a matrix showing the model's score for each image against each possible choice category.

Model output for 2AFC: the final model output saved to the CSV is accuracy, one value per image.

Step 5: Save results

Object 2AFC saves:

df = pd.DataFrame({
    "image_name": images,
    "image_index": indices,
    "accuracy": image_scores
})

df.to_csv("results/obj_2afc.csv", index=False)

Drawing 2AFC saves the same kind of table:

df.to_csv("results/draw_2afc.csv", index=False)

7. Ratings notebook

The notebook ratings.ipynb tests a face-rating model on the images from the ratings task.

What goes in?

The notebook reads face images from:

../images/ratings/

These files are .jpg images.

indices = sort_images(os.listdir("../images/ratings/"))

image_path = f"../images/ratings/im{indices[img_idx]}.jpg"

What model is used?

The notebook uses the ComboNet model, a model to judge facial aesthetics, which we will find and download from GitHub first. The model weights are loaded from:

checkpoints/ComboNet_SCUTFBP5500.pth
fbp = FacialPredictor(
    pretrained_model_path="checkpoints/ComboNet_SCUTFBP5500.pth"
)

What does the model output?

For each face image, the model outputs one numerical rating score (between 1 and 5, where 1 = lowest "aesthetics" score and 5 = highest "aesthetics" score).

ratings = []

for img_idx in range(len(indices)):
    image_path = f"../images/ratings/im{indices[img_idx]}.jpg"
    rating = fbp.infer(image_path)
    ratings.append(rating)

The function fbp.infer loads one image, prepares it for the model, runs the model, and returns one number.

Model output for ratings: the final model output saved to the CSV is score, one rating score per image.

What gets saved?

df = pd.DataFrame({
    "image_name": images,
    "image_index": indices,
    "score": np.array(ratings)
})

df.to_csv("results/ratings.csv", index=False)

8. N-back notebook

The notebook n_back.ipynb tests a memorability model on the images used in the n-back task.

What goes in?

The notebook reads images from:

../images/n_back/

These files are .png images.

indices = sort_images(os.listdir("../images/n_back/"))

image_path = f"../images/n_back/im{indices[img_idx]}.png"

What model is used?

The notebook uses ViTMem, a model that predicts image memorability.

from vitmem import ViTMem

vitmem = ViTMem()

What does the model output?

For each image, the model outputs one memorability score. This score can be compared with behavior from the memory-related task.

scores_vitmem = []

for img_idx in range(len(indices)):
    image_path = f"../images/n_back/im{indices[img_idx]}.png"
    memorability = vitmem(image_path)
    scores_vitmem.append(memorability)
Model output for n-back: the final model output saved to the CSV is score, one memorability score per image.

What gets saved?

df = pd.DataFrame({
    "image_name": images,
    "image_index": indices,
    "score": np.array(scores_vitmem)
})

df.to_csv("results/n_back.csv", index=False)

9. Output files

Every simplified notebook saves one CSV file. Each row corresponds to one image.

Notebook Output CSV Columns Meaning of model output
obj_2afc.ipynb results/obj_2afc.csv image_name, image_index, accuracy 2AFC-style object recognition accuracy for each object image.
draw_2afc.ipynb results/draw_2afc.csv image_name, image_index, accuracy 2AFC-style object recognition accuracy for each drawing image.
ratings.ipynb results/ratings.csv image_name, image_index, score Predicted face rating score for each image.
n_back.ipynb results/n_back.csv image_name, image_index, score Predicted memorability score for each image.
Most important columns: image_name and image_index. These keep the link between each image, the model output, and the human data.

10. Summary

Main takeaway: the point is not only to test whether the model performs well. The point is to test whether the model behaves like a human subject on the same task, and to find where that model-human similarity breaks down.

11. Student checklist

Use this checklist when you want to test a model like a subject. You can change the model, change the image set, or test altered images. The important thing is to keep the task structure and output format clear.

1. Choose the task

Start by choosing the notebook that matches the task you want to test.

Task Notebook Model output
Object 2AFC obj_2afc.ipynb accuracy
Drawing 2AFC draw_2afc.ipynb accuracy
Ratings ratings.ipynb score
N-back / memorability n_back.ipynb score

2. Choose or change the model

You can test different models. For example, in the 2AFC notebooks, the model currently uses AlexNet:

model = models.alexnet(pretrained=True)

Students can replace it with another model, such as ResNet50:

model = models.resnet50(pretrained=True)

Or they can test another architecture, such as ConvNeXt:

model = timm.create_model("convnext_base", pretrained=True)

After changing the model, make sure the model is in evaluation mode:

model.eval()
Check carefully: if you change the model, the output layer may also need to be adjusted so the model predicts the correct number of task categories.

3. Make sure the model output matches the task

Different tasks need different outputs.

Ask yourself: Does my model output one value that makes sense for this task?

4. Choose or change the image folder

You can test the original task images or a new image set. For example:

data_loader = make_val_loader("../images/obj_2AFC", 8)

can be changed to a folder containing altered images:

data_loader = make_val_loader("../images/obj_2AFC_blur", 8)

or:

data_loader = make_val_loader("../images/obj_2AFC_noise", 8)

This is useful for testing whether the model is robust to image changes.

5. Keep the same image naming format

The code expects image names to contain image numbers. For most tasks, images should be named like this:

im1.png
im2.png
im3.png
...

For ratings, images may be named like this:

im1.jpg
im2.jpg
im3.jpg
...

The image number is important because it keeps the model output aligned with the human data.

6. Check the labels for 2AFC tasks

For 2AFC tasks, the labels tell the code which category is correct for each image. Example:

labels = [0, 0, 1, 1, 2, 2, 2, 3, 4, 4, 5, 6, 6, 7, 7, 8]

Before running a new image set, check that the number of labels matches the number of images:

len(labels)

Also check that each label corresponds to the correct image.

7. Test altered images

You can create altered versions of the images and test the model again. Examples include:

Scientific question: Does the model still behave like humans when the images are altered?

8. Run the notebook from top to bottom

After changing the model or images, run all cells from the beginning. Then check that the notebook produced one output per image.

For 2AFC tasks:

len(image_scores)

For ratings:

len(ratings)

For n-back:

len(scores_vitmem)

9. Save results with a clear file name

Avoid overwriting previous results. Instead of always saving:

df.to_csv("results/obj_2afc.csv", index=False)

use a more specific file name:

df.to_csv("results/obj_2afc_alexnet_original.csv", index=False)

or:

df.to_csv("results/obj_2afc_alexnet_blur.csv", index=False)

Good file names include the task, model, and image condition.

obj_2afc_alexnet_original.csv
obj_2afc_resnet50_original.csv
obj_2afc_alexnet_blur.csv
draw_2afc_alexnet_noise.csv
ratings_combonet_original.csv
n_back_vitmem_contrast.csv

10. Check the output CSV

Every output file should have one row per image.

For 2AFC tasks:

image_name, image_index, accuracy

For ratings and n-back:

image_name, image_index, score

You can quickly check the table with:

df.head()

Final checklist before saving

[ ] I chose the correct notebook.

[ ] I know what task the model is doing.

[ ] I know what the model output means.

[ ] I used the correct image folder.

[ ] My images are named correctly.

[ ] My labels match my images.

[ ] I changed the output CSV name if needed.

[ ] The CSV has one row per image.

[ ] The CSV keeps image_name and image_index.

[ ] I can compare the model output with human behavior.

12. How do we validate the models?

After the notebooks save model outputs, we compare the model results with human behavior.

Question 1: Does the model do the task?

For 2AFC tasks, we check whether the model has high accuracy. For rating and n-back tasks, we check whether the model gives reasonable scores.

Question 2: Where does the model break?

The failures are important. If models fail on specific images, this tells us what the model does not capture about human behavior.

Question 3:

When you manipulate the images, what do you expect will happen to task performance for humans and models? Will they perform the same or different? What can you conclude from this experiment if your hypothesis is confirmed?

13. Summary

Main takeaway 1: the point is not only to test whether the model performs well. The point is to test whether the model behaves like a human subject on the same task, and to find where that model-human similarity breaks down.
Main takeaway 2: By using models that "act" similar to humans (i.e., they take the same inputs as humans and produce the same outputs), we can create falsifiable predictions for new images or tasks the model has not seen before.
Hub