Simulation only. No human participants have been recruited or run. Every figure on this site is computed from synthetic observers with known ground truth, and says so.

Red team

Every observer on this page has, by construction, no stimulus-history effect at all. The only acceptable answer from the analysis is “nothing here”.

Five generators were used to try to manufacture serial dependence: an observer with no history; one that regresses toward the running mean of recent orientations; one attracted to cardinal orientations; one whose dial drifts toward its random start angle; and one biased toward its own previous response. Each was run 200 times on the balanced stimulus design and again on a deliberately autocorrelated one. Each simulated “experiment” resamples 24 of them and applies the primary test.

The naive lag-1 model, which is what most of the literature fits, reports serial dependence for the central-tendency observer in 100% of experiments. The adjusted primary model reports it in 5.5%.

Fig. 1 False-positive rate by generator and model, balanced design

SIMULATED DATA

Each row: proportion of simulated 24-observer experiments in which the observer-level test rejected a_stim = 0 at α = 0.05. The vertical rule is the nominal 5%.Source: results/red_team.parquet · run red_team-e756953-4000
Table view (15 rows)
generatormodelfalse-positive ratecalibrated ratebias (°)bias SEbias pgate
cardinal attractionnaive lag-1 (M1)5.5%5.6%−0.0510.0700.466
cardinal attractionstim + response (M3)33.0%5.5%−0.3220.071< 0.001
cardinal attractionadjusted primary (M3adj)9.2%5.5%−0.0970.0610.114pass
central tendencynaive lag-1 (M1)100.0%4.8%+1.3100.043< 0.001
central tendencystim + response (M3)100.0%5.7%+1.3480.039< 0.001
central tendencyadjusted primary (M3adj)5.5%4.8%−0.0240.0620.702pass
motor persistencenaive lag-1 (M1)5.9%4.5%+0.0640.0880.468
motor persistencestim + response (M3)4.6%4.6%+0.0270.0920.774
motor persistenceadjusted primary (M3adj)9.3%4.7%−0.1060.0590.072pass
null (no history)naive lag-1 (M1)9.7%5.1%−0.1140.0590.054
null (no history)stim + response (M3)7.0%4.6%−0.0650.0630.304
null (no history)adjusted primary (M3adj)8.8%4.9%−0.1030.0600.086pass
response historynaive lag-1 (M1)99.9%6.9%+1.0860.065< 0.001
response historystim + response (M3)13.6%5.0%+0.1580.0610.010
response historyadjusted primary (M3adj)13.6%5.6%+0.1590.0610.009FAIL

Two different ways to fail

A rejection rate estimated by resampling from a finite pool of observers mixes two things. One is genuine estimator bias: the confound really does shift the amplitude estimate. The other is how the test behaves given the shape of the estimator’s sampling distribution. The table separates them: the bias column is the pool mean with its standard error and a t-test against zero; the calibrated rate resamples from the pool re-centred at zero. Hypothesis H6 passes only if the calibrated rate is at most 7.5% and the bias is not significant.

The adjusted model does not pass for every generator. It fails the gate for response history (bias +0.16° ± 0.06, p 0.009). The response-history leak has a known cause: the synthetic observer is biased toward its previous response relative to its own noisy percept, but the model can only condition on the observable target. That errors-in-variables gap is not closable with behavioural data alone, and it is reported as a limitation rather than tuned away.

Fig. 2 The same test on an autocorrelated stimulus sequence

SIMULATED DATA

The same generators when consecutive orientations are correlated (a random walk). This is the design TimeMind does not use; the sequence diagnostic in the pipeline exists to catch it.Source: results/red_team.parquet · run red_team-e756953-4000
Table view (15 rows)
generatormodelfalse-positive ratebias (°)bias p
cardinal attractionnaive lag-1 (M1)5.8%+0.0300.552
cardinal attractionstim + response (M3)65.5%−0.386< 0.001
cardinal attractionadjusted primary (M3adj)4.9%−0.0010.988
central tendencynaive lag-1 (M1)100.0%+1.655< 0.001
central tendencystim + response (M3)100.0%+1.347< 0.001
central tendencyadjusted primary (M3adj)5.1%+0.0080.867
motor persistencenaive lag-1 (M1)6.3%−0.0730.280
motor persistencestim + response (M3)5.7%−0.0600.431
motor persistenceadjusted primary (M3adj)4.7%+0.0200.669
null (no history)naive lag-1 (M1)4.9%+0.0280.526
null (no history)stim + response (M3)5.9%+0.0520.279
null (no history)adjusted primary (M3adj)10.9%+0.1030.047
response historynaive lag-1 (M1)100.0%+1.229< 0.001
response historystim + response (M3)9.4%+0.0990.054
response historyadjusted primary (M3adj)19.1%+0.1630.002

Can model comparison find the true generator?

The complement of a false-positive test is a discrimination test: generate from a known model, fit them all, and see which one wins. Picking exactly the right model is hard — AIC does it in 81% of simulations and a single held-out split in 57%. The question that matters for a claim about perception is coarser: did the winning model contain a stimulus-history term when, and only when, the generator had one? AIC gets that verdict right in 91% of simulations overall, and held-out likelihood in 72%.

Fig. 3 Which model wins, by generating observer

SIMULATED DATA

Rows: the observer that generated the data. Columns: the model with the highest held-out log-likelihood. Outlined cells are the correct answer. Each row sums to 100%.Source: results/model_confusion.parquet · run model_confusion-e756953-3000
Table view (8 rows)
generatorcorrect stim-history verdict (AIC)
null (no history)85%
attractive (true effect)100%
repulsive (true effect)100%
response history95%
central tendency90%
cardinal attraction78%
motor persistence78%
multi-lag (true effect)100%