Red team
Every observer on this page has, by construction, no stimulus-history effect at all. The only acceptable answer from the analysis is “nothing here”.
Five generators were used to try to manufacture serial dependence: an observer with no history; one that regresses toward the running mean of recent orientations; one attracted to cardinal orientations; one whose dial drifts toward its random start angle; and one biased toward its own previous response. Each was run 200 times on the balanced stimulus design and again on a deliberately autocorrelated one. Each simulated “experiment” resamples 24 of them and applies the primary test.
The naive lag-1 model, which is what most of the literature fits, reports serial dependence for the central-tendency observer in 100% of experiments. The adjusted primary model reports it in 5.5%.
Fig. 1 False-positive rate by generator and model, balanced design
SIMULATED DATA
false-positive rate · vertical rule: nominal α = 0.05
results/red_team.parquet · run red_team-e756953-4000Table view (15 rows)
| generator | model | false-positive rate | calibrated rate | bias (°) | bias SE | bias p | gate |
|---|---|---|---|---|---|---|---|
| cardinal attraction | naive lag-1 (M1) | 5.5% | 5.6% | −0.051 | 0.070 | 0.466 | — |
| cardinal attraction | stim + response (M3) | 33.0% | 5.5% | −0.322 | 0.071 | < 0.001 | — |
| cardinal attraction | adjusted primary (M3adj) | 9.2% | 5.5% | −0.097 | 0.061 | 0.114 | pass |
| central tendency | naive lag-1 (M1) | 100.0% | 4.8% | +1.310 | 0.043 | < 0.001 | — |
| central tendency | stim + response (M3) | 100.0% | 5.7% | +1.348 | 0.039 | < 0.001 | — |
| central tendency | adjusted primary (M3adj) | 5.5% | 4.8% | −0.024 | 0.062 | 0.702 | pass |
| motor persistence | naive lag-1 (M1) | 5.9% | 4.5% | +0.064 | 0.088 | 0.468 | — |
| motor persistence | stim + response (M3) | 4.6% | 4.6% | +0.027 | 0.092 | 0.774 | — |
| motor persistence | adjusted primary (M3adj) | 9.3% | 4.7% | −0.106 | 0.059 | 0.072 | pass |
| null (no history) | naive lag-1 (M1) | 9.7% | 5.1% | −0.114 | 0.059 | 0.054 | — |
| null (no history) | stim + response (M3) | 7.0% | 4.6% | −0.065 | 0.063 | 0.304 | — |
| null (no history) | adjusted primary (M3adj) | 8.8% | 4.9% | −0.103 | 0.060 | 0.086 | pass |
| response history | naive lag-1 (M1) | 99.9% | 6.9% | +1.086 | 0.065 | < 0.001 | — |
| response history | stim + response (M3) | 13.6% | 5.0% | +0.158 | 0.061 | 0.010 | — |
| response history | adjusted primary (M3adj) | 13.6% | 5.6% | +0.159 | 0.061 | 0.009 | FAIL |
Two different ways to fail
A rejection rate estimated by resampling from a finite pool of observers mixes two things. One is genuine estimator bias: the confound really does shift the amplitude estimate. The other is how the test behaves given the shape of the estimator’s sampling distribution. The table separates them: the bias column is the pool mean with its standard error and a t-test against zero; the calibrated rate resamples from the pool re-centred at zero. Hypothesis H6 passes only if the calibrated rate is at most 7.5% and the bias is not significant.
The adjusted model does not pass for every generator. It fails the gate for response history (bias +0.16° ± 0.06, p 0.009). The response-history leak has a known cause: the synthetic observer is biased toward its previous response relative to its own noisy percept, but the model can only condition on the observable target. That errors-in-variables gap is not closable with behavioural data alone, and it is reported as a limitation rather than tuned away.
Fig. 2 The same test on an autocorrelated stimulus sequence
SIMULATED DATA
false-positive rate · vertical rule: nominal α = 0.05
results/red_team.parquet · run red_team-e756953-4000Table view (15 rows)
| generator | model | false-positive rate | bias (°) | bias p |
|---|---|---|---|---|
| cardinal attraction | naive lag-1 (M1) | 5.8% | +0.030 | 0.552 |
| cardinal attraction | stim + response (M3) | 65.5% | −0.386 | < 0.001 |
| cardinal attraction | adjusted primary (M3adj) | 4.9% | −0.001 | 0.988 |
| central tendency | naive lag-1 (M1) | 100.0% | +1.655 | < 0.001 |
| central tendency | stim + response (M3) | 100.0% | +1.347 | < 0.001 |
| central tendency | adjusted primary (M3adj) | 5.1% | +0.008 | 0.867 |
| motor persistence | naive lag-1 (M1) | 6.3% | −0.073 | 0.280 |
| motor persistence | stim + response (M3) | 5.7% | −0.060 | 0.431 |
| motor persistence | adjusted primary (M3adj) | 4.7% | +0.020 | 0.669 |
| null (no history) | naive lag-1 (M1) | 4.9% | +0.028 | 0.526 |
| null (no history) | stim + response (M3) | 5.9% | +0.052 | 0.279 |
| null (no history) | adjusted primary (M3adj) | 10.9% | +0.103 | 0.047 |
| response history | naive lag-1 (M1) | 100.0% | +1.229 | < 0.001 |
| response history | stim + response (M3) | 9.4% | +0.099 | 0.054 |
| response history | adjusted primary (M3adj) | 19.1% | +0.163 | 0.002 |
Can model comparison find the true generator?
The complement of a false-positive test is a discrimination test: generate from a known model, fit them all, and see which one wins. Picking exactly the right model is hard — AIC does it in 81% of simulations and a single held-out split in 57%. The question that matters for a claim about perception is coarser: did the winning model contain a stimulus-history term when, and only when, the generator had one? AIC gets that verdict right in 91% of simulations overall, and held-out likelihood in 72%.
Fig. 3 Which model wins, by generating observer
SIMULATED DATA
selected model
generating observer (rows) · cells show % of simulations
results/model_confusion.parquet · run model_confusion-e756953-3000Table view (8 rows)
| generator | correct stim-history verdict (AIC) |
|---|---|
| null (no history) | 85% |
| attractive (true effect) | 100% |
| repulsive (true effect) | 100% |
| response history | 95% |
| central tendency | 90% |
| cardinal attraction | 78% |
| motor persistence | 78% |
| multi-lag (true effect) | 100% |