ANN notebooks
Use these launch buttons to open the original notebooks in Google Colab. The tutorial content below is the original version. For each demo, make a copy of these original files to your Drive.
anns/AlexNet_demo.ipynbanns/Facial_emotion.ipynbanns/obj_2afc.ipynbanns/draw_2afc.ipynbanns/ratings.ipynbanns/n_back.ipynbcompare_humans_vs_anns.ipynb1. Goal of this tutorial
This tutorial explains how to make the models perform the same task as humans. The main idea is very simple:
In the psychophysics experiments, a human participant saw images and gave responses. In the ANN notebooks, a model will see the same images and perform the same task as the human participants, producing the similar kind of responses. We call this the "model output".
2. What does it mean to "treat a model as a human subject"?
A human subject sees an image and produces a response. A model also receives an image and produces an output. The code makes these two things comparable.
The model output depends on the task. In a recognition task, the model may output probabilities for object categories. In a rating task, the model outputs a score. In an n-back task, the model outputs a memorability score for each image.
3. Demo notebooks
Before running model experiments with "task-matched" notebooks, we will learn what happens in these experiments using demo notebooks. These demo notebooks are simpler than the task-matched notebooks, and they show how to load a pretrained model, how give it one image, and how read the model prediction.
Why do we include demo notebooks?
The demo notebooks are useful because they introduce the basic model workflow:
- load a pretrained model,
- prepare an image so the model can read it,
- run the image through the model,
- look at the model prediction.
| Notebook | What the model was trained to do | What goes in | What comes out |
|---|---|---|---|
AlexNet_demo.ipynb |
AlexNet was trained for image classification. It learns to look at an image and predict what object category is present. | One image. | A list of predicted object labels, usually ordered from most likely to least likely. |
Facial_emotion.ipynb |
The facial emotion model was trained to classify facial expressions. It learns to look at a face and predict an emotion label. | One face image. | A list of predicted emotion labels and confidence scores. |
AlexNet demo
In the AlexNet demo, the notebook loads a pretrained image-classification model, prepares one image, runs the image through the model, and produces the 5 most likely object categories.
The output is not a human 2AFC response yet. It is simply the model saying which object categories it thinks are most likely.
model = models.alexnet(pretrained=True)
model.eval()
with torch.no_grad():
output = model(input_image)
_, indices = torch.topk(output, 5)
In simple terms, this asks: What object does the pretrained model think is in this image?
Facial emotion demo
In the facial emotion demo, a pretrained face-expression model is used. The notebook gives the model one face image and produces the most likely emotion label.
clf = pipeline(
"image-classification",
model="trpakov/vit-face-expression"
)
results = clf(img)
In simple terms, this asks: What emotion does the pretrained model think this face shows?
After understanding this basic idea, the task-matched notebooks extend the same logic to the full psychophysics tasks.
4. Task-matched notebook guide
| Notebook | Human task | What the model does | Output file |
|---|---|---|---|
obj_2afc.ipynb |
Object 2AFC: humans choose the correct object from two options. | A recognition model predicts object probabilities. The code converts those probabilities into 2AFC-style image accuracy. | results/obj_2afc.csv |
draw_2afc.ipynb |
Drawing 2AFC: humans choose the correct object from two options, but the image is a drawing or sketch. | The same recognition model is tested on drawings. The code again converts probabilities into 2AFC-style image accuracy. | results/draw_2afc.csv |
ratings.ipynb |
Ratings: humans rate face images. | A face-rating model predicts one score for each face image. | results/ratings.csv |
n_back.ipynb |
N-back: humans perform a memory-related task with images. | A memorability model predicts one memorability score for each image. | results/n_back.csv |
6. Object and drawing 2AFC notebooks
The notebooks obj_2afc.ipynb and draw_2afc.ipynb
use the same logic. The difference is the image folder:
obj_2afc.ipynbuses../images/obj_2AFCdraw_2afc.ipynbuses../images/draw_2AFC
What is a 2AFC task?
2AFC means two-alternative forced choice. The participant chooses between two possible answers.
For example:
The model does not choose by clicking a button. Instead, it gives a probability for each object class. The code uses these probabilities to ask which option the model would prefer.
Model used in both 2AFC notebooks
The notebooks currently load AlexNet.
AlexNet starts as a pretrained image-classification model.
Then the final layer is changed so the model predicts 10 task categories.
Finally, fine-tuned weights are loaded from checkpoints/alexnet_fine_tuned.pth.
model = models.alexnet(pretrained=True)
weights = torch.load("checkpoints/alexnet_fine_tuned.pth")
model.classifier[6] = nn.Linear(model.classifier[6].in_features, 10)
model.load_state_dict(weights)
model.eval()
The notebooks also include commented code for other possible models,
such as ResNet50 and ConvNeXt.
Step 1: Load task images and labels
For object 2AFC:
data_loader = make_val_loader("../images/obj_2AFC", 8)
labels = [0, 0, 1, 1, 2, 2, 2, 3, 4, 4, 5, 6, 6, 7, 7, 8]
For drawing 2AFC:
data_loader = make_val_loader("../images/draw_2AFC", 4)
labels = [0, 1, 4, 5]
The labels tell the code which object category is correct for each image. The labels used in these tasks are defined in the metadata README: README_metadata_tests.md .
Step 2: Get model probabilities
The model is run on every image. Its output is in the form of probabilities per object category class.
probabilities = []
with torch.no_grad():
for inputs in data_loader:
inputs = inputs[0]
output = model(inputs)
output = torch.nn.functional.softmax(output, dim=-1)
probabilities.append(output.detach().cpu().numpy())
probabilities = np.concatenate(probabilities, axis=0)
After this step, each image has 10 probabilities: one probability for each possible object category.
Step 3: Check normal classification accuracy
The notebook first checks the usual model accuracy. This asks whether the model's highest-probability category is the correct category.
preds = np.argmax(probabilities, axis=1)
(preds == labels).mean()
This is useful, but it is not exactly the human 2AFC task. Humans choose between two options, not all 10 categories at once.
Step 4: Convert probabilities into 2AFC-like accuracy
This is the key step. The code compares the model probability for the correct answer with the probability for each possible wrong answer.
For each image, the model checks how likely it would have chosen the correct category label compared to picking one of the wrong category labels:
score =
probability_correct_category /
(probability_correct_category + probability_wrong_category)
This gives a model 2AFC score. If the model gives more probability to the correct answer than to the wrong answer, the score is high. If it gives more probability to the wrong answer, the score is low. The score is computed across all possible distractors and averaged to produce a single score per image (called image_scores below).
image_scores, model_choices = create_image_level_accuracies(
probabilities,
labels=labels
)
The output image_scores gives one 2AFC-style accuracy score per image.
The output model_choices is a matrix showing the model's score for each image
against each possible choice category.
accuracy,
one value per image.
Step 5: Save results
Object 2AFC saves:
df = pd.DataFrame({
"image_name": images,
"image_index": indices,
"accuracy": image_scores
})
df.to_csv("results/obj_2afc.csv", index=False)
Drawing 2AFC saves the same kind of table:
df.to_csv("results/draw_2afc.csv", index=False)
7. Ratings notebook
The notebook ratings.ipynb tests a face-rating model on the images
from the ratings task.
What goes in?
The notebook reads face images from:
../images/ratings/
These files are .jpg images.
indices = sort_images(os.listdir("../images/ratings/"))
image_path = f"../images/ratings/im{indices[img_idx]}.jpg"
What model is used?
The notebook uses the ComboNet model, a model to judge facial aesthetics, which we will find and download from GitHub first. The model weights are loaded from:
checkpoints/ComboNet_SCUTFBP5500.pth
fbp = FacialPredictor(
pretrained_model_path="checkpoints/ComboNet_SCUTFBP5500.pth"
)
What does the model output?
For each face image, the model outputs one numerical rating score (between 1 and 5, where 1 = lowest "aesthetics" score and 5 = highest "aesthetics" score).
ratings = []
for img_idx in range(len(indices)):
image_path = f"../images/ratings/im{indices[img_idx]}.jpg"
rating = fbp.infer(image_path)
ratings.append(rating)
The function fbp.infer loads one image, prepares it for the model,
runs the model, and returns one number.
score,
one rating score per image.
What gets saved?
df = pd.DataFrame({
"image_name": images,
"image_index": indices,
"score": np.array(ratings)
})
df.to_csv("results/ratings.csv", index=False)
8. N-back notebook
The notebook n_back.ipynb tests a memorability model on the images
used in the n-back task.
What goes in?
The notebook reads images from:
../images/n_back/
These files are .png images.
indices = sort_images(os.listdir("../images/n_back/"))
image_path = f"../images/n_back/im{indices[img_idx]}.png"
What model is used?
The notebook uses ViTMem, a model that predicts image memorability.
from vitmem import ViTMem
vitmem = ViTMem()
What does the model output?
For each image, the model outputs one memorability score. This score can be compared with behavior from the memory-related task.
scores_vitmem = []
for img_idx in range(len(indices)):
image_path = f"../images/n_back/im{indices[img_idx]}.png"
memorability = vitmem(image_path)
scores_vitmem.append(memorability)
score,
one memorability score per image.
What gets saved?
df = pd.DataFrame({
"image_name": images,
"image_index": indices,
"score": np.array(scores_vitmem)
})
df.to_csv("results/n_back.csv", index=False)
9. Output files
Every simplified notebook saves one CSV file. Each row corresponds to one image.
| Notebook | Output CSV | Columns | Meaning of model output |
|---|---|---|---|
obj_2afc.ipynb |
results/obj_2afc.csv |
image_name, image_index, accuracy |
2AFC-style object recognition accuracy for each object image. |
draw_2afc.ipynb |
results/draw_2afc.csv |
image_name, image_index, accuracy |
2AFC-style object recognition accuracy for each drawing image. |
ratings.ipynb |
results/ratings.csv |
image_name, image_index, score |
Predicted face rating score for each image. |
n_back.ipynb |
results/n_back.csv |
image_name, image_index, score |
Predicted memorability score for each image. |
image_name and image_index.
These keep the link between each image, the model output, and the human data.
10. Summary
- The goal is to test models like subjects.
- Each notebook loads the task images.
- The images are sorted so the model output stays linked to the correct image.
- The model produces one output per image.
- For
obj_2afc.ipynb, the output isaccuracy. - For
draw_2afc.ipynb, the output isaccuracy. - For
ratings.ipynb, the output isscore. - For
n_back.ipynb, the output isscore. - Each notebook saves a CSV file with
image_nameandimage_index. - Those columns allow model behavior and human behavior to be compared image by image.
11. Student checklist
Use this checklist when you want to test a model like a subject. You can change the model, change the image set, or test altered images. The important thing is to keep the task structure and output format clear.
1. Choose the task
Start by choosing the notebook that matches the task you want to test.
| Task | Notebook | Model output |
|---|---|---|
| Object 2AFC | obj_2afc.ipynb |
accuracy |
| Drawing 2AFC | draw_2afc.ipynb |
accuracy |
| Ratings | ratings.ipynb |
score |
| N-back / memorability | n_back.ipynb |
score |
2. Choose or change the model
You can test different models. For example, in the 2AFC notebooks, the model currently uses AlexNet:
model = models.alexnet(pretrained=True)
Students can replace it with another model, such as ResNet50:
model = models.resnet50(pretrained=True)
Or they can test another architecture, such as ConvNeXt:
model = timm.create_model("convnext_base", pretrained=True)
After changing the model, make sure the model is in evaluation mode:
model.eval()
3. Make sure the model output matches the task
Different tasks need different outputs.
-
For 2AFC tasks, the model should output probabilities over object classes.
The code converts these probabilities into 2AFC-style
accuracy. -
For ratings, the model should output one rating
scoreper image. -
For n-back, the model should output one memorability
scoreper image.
4. Choose or change the image folder
You can test the original task images or a new image set. For example:
data_loader = make_val_loader("../images/obj_2AFC", 8)
can be changed to a folder containing altered images:
data_loader = make_val_loader("../images/obj_2AFC_blur", 8)
or:
data_loader = make_val_loader("../images/obj_2AFC_noise", 8)
This is useful for testing whether the model is robust to image changes.
5. Keep the same image naming format
The code expects image names to contain image numbers. For most tasks, images should be named like this:
im1.png
im2.png
im3.png
...
For ratings, images may be named like this:
im1.jpg
im2.jpg
im3.jpg
...
The image number is important because it keeps the model output aligned with the human data.
6. Check the labels for 2AFC tasks
For 2AFC tasks, the labels tell the code which category is correct for each image. Example:
labels = [0, 0, 1, 1, 2, 2, 2, 3, 4, 4, 5, 6, 6, 7, 7, 8]
Before running a new image set, check that the number of labels matches the number of images:
len(labels)
Also check that each label corresponds to the correct image.
7. Test altered images
You can create altered versions of the images and test the model again. Examples include:
- blurred images,
- noisy images,
- low-contrast images,
- cropped images,
- rotated images,
- grayscale images.
8. Run the notebook from top to bottom
After changing the model or images, run all cells from the beginning. Then check that the notebook produced one output per image.
For 2AFC tasks:
len(image_scores)
For ratings:
len(ratings)
For n-back:
len(scores_vitmem)
9. Save results with a clear file name
Avoid overwriting previous results. Instead of always saving:
df.to_csv("results/obj_2afc.csv", index=False)
use a more specific file name:
df.to_csv("results/obj_2afc_alexnet_original.csv", index=False)
or:
df.to_csv("results/obj_2afc_alexnet_blur.csv", index=False)
Good file names include the task, model, and image condition.
obj_2afc_alexnet_original.csv
obj_2afc_resnet50_original.csv
obj_2afc_alexnet_blur.csv
draw_2afc_alexnet_noise.csv
ratings_combonet_original.csv
n_back_vitmem_contrast.csv
10. Check the output CSV
Every output file should have one row per image.
For 2AFC tasks:
image_name, image_index, accuracy
For ratings and n-back:
image_name, image_index, score
You can quickly check the table with:
df.head()
Final checklist before saving
[ ] I chose the correct notebook.
[ ] I know what task the model is doing.
[ ] I know what the model output means.
[ ] I used the correct image folder.
[ ] My images are named correctly.
[ ] My labels match my images.
[ ] I changed the output CSV name if needed.
[ ] The CSV has one row per image.
[ ] The CSV keeps image_name and image_index.
[ ] I can compare the model output with human behavior.
12. How do we validate the models?
After the notebooks save model outputs, we compare the model results with human behavior.
Question 1: Does the model do the task?
For 2AFC tasks, we check whether the model has high accuracy. For rating and n-back tasks, we check whether the model gives reasonable scores.
Question 2: Where does the model break?
The failures are important. If models fail on specific images, this tells us what the model does not capture about human behavior.
Question 3:
When you manipulate the images, what do you expect will happen to task performance for humans and models? Will they perform the same or different? What can you conclude from this experiment if your hypothesis is confirmed?13. Summary
- The goal is to test models like subjects.
- Each notebook loads the task images.
- The images are sorted so the model output stays linked to the correct image.
- The model produces one output per image.
- For
obj_2afc.ipynb, the output isaccuracy. - For
draw_2afc.ipynb, the output isaccuracy. - For
ratings.ipynb, the output isscore. - For
n_back.ipynb, the output isscore. - Each notebook saves a CSV file with
image_nameandimage_index. - Those columns allow model behavior and human behavior to be compared image by image.