Summer School 2026
Class project

From classroom behavior to reliability, noise correction, and interpretation.

In this project, students use the performance data collected from the summer school tasks and ask a deeper question: how reliable are the measured behavioral patterns, and how should we interpret correlations when measurements are noisy?

Project overview

This activity turns the tutorial into a mini research project using the class's own behavioral data.

1. Run the tasks

Students complete the browser experiments and save the downloaded CSV files.

2. Pool performance

The class-level CSV files are combined across participants for each task.

3. Estimate reliability

Repeated measurements are split into halves to ask whether the same images or conditions produce consistent patterns.

4. Learn noise correction

A separate simulation notebook shows how measurement noise lowers observed correlations and how reliability can help interpret them.

5. Interpret the project

Students discuss which behavioral measurements are stable, which are noisy, and what that means for comparing humans and models.

Core idea

A low correlation does not always mean that two systems are unrelated. Sometimes the measurement itself is noisy.

Noise tutorial

Why correlations can shrink

The noise correction notebook starts with variables whose true relationship is known, then adds different amounts of measurement noise. Students can see that the observed correlation drops as the measurement gets noisier.

This prepares students to interpret real behavioral and model comparisons more carefully.

Task reliability

How consistent is the class data?

The split-half reliability notebook uses the actual task data to ask whether different random halves of the class give similar image-level or condition-level patterns.

If the split halves agree, the measurement is stable. If they disagree, we should be cautious about over-interpreting the result.

Class behavioral datasets

These CSV files contain the aggregate student performance used for the reliability activity.

2AFC

Object recognition

Class responses from the object two-alternative forced-choice task.

2AFC

Drawing recognition

Class responses from the drawing two-alternative forced-choice task.

Ratings

Image ratings

Class image-rating responses used to estimate image-level consistency.

Memory

N-back

Class responses from the image memory / repetition task.

Tip: download the CSV files or access them from the shared Drive folder before running the reliability notebook. If you are using Colab, first add the CVR tutorial folder to your own My Drive.

Project notebooks

Open these in Google Colab. Save a copy to your own Drive before editing.

Concept notebook

Noise correction tutorial

Use simulations to see how measurement noise weakens observed correlations, how split-half reliability is estimated, and how reliability can be used to correct correlations.

Class data notebook

Reliability from task data

Use the class performance CSV files to estimate split-half reliability for the behavioral tasks and ask which measurements are most stable.

NotebookOpenUse it for
Noise correction tutorialOpen in ColabUnderstand how noise reduces observed correlations and how reliability estimates help interpret them.
Reliability from task dataOpen in ColabEstimate split-half reliability from the class behavioral datasets.

Discussion questions

Use these prompts for a short project presentation or group discussion.

Reliability

Which task produced the most reliable image-level or condition-level pattern? Which task was noisiest? What might explain the difference?

Number of repetitions

How would reliability change if each task had more participants, more repeats, or fewer stimuli?

Human/model comparison

If a human-vs-ANN correlation is low, how can we tell whether the model is poor or whether the behavioral measurement is noisy?

Experimental design

What would you change in the task design to make the measurements more reliable?