Simulation only. No human participants have been recruited or run. Every figure on this site is computed from synthetic observers with known ground truth, and says so.

Past Shapes Present: Distinguishing Stimulus History, Response History, and Central Tendency in Serial Dependence — a browser platform and estimator benchmark

TimeMind manuscript · generated 2026-09-11T20:30+00:00 from results at git `e756953` · study status `SIMULATION_ONLY`

SIMULATED DATA. STUDY STATUS: SIMULATION_ONLY. No human participants have been recruited, consented, or run. Every quantitative statement below describes the behaviour of an estimator on synthetic observers with known ground truth. No claim about human perception is made.

Abstract

Serial dependence — the pull of a current perceptual judgement toward recent stimuli — is well replicated in orientation estimation, but a raw lag-1 correlation cannot distinguish a perceptual history effect from attraction to one's own previous response, regression toward the running mean of recent stimuli, attraction to cardinal orientations, persistence at the response device, or temporal autocorrelation in the stimulus sequence. We present TimeMind, a reproducible browser experiment for continuous orientation estimation with a post-stimulus orthogonal-report cue (after Cicchini, Mikellidou & Burr, 2017), a frozen exclusion policy, and an analysis pipeline whose primary model fits stimulus-history and response-history tuning curves jointly with the nuisance processes above. We validate the pipeline with synthetic observers of known ground truth. The lag-1 amplitude is recovered with worst-case absolute bias 0.22° at 600 trials per observer; stimulus- and response-history amplitudes are jointly identifiable (recovery-error correlation r = −0.40) with 50% orthogonal trials but not without them (r = −0.88). A central-tendency observer with no serial dependence is declared serially dependent by the standard lag-1 model in 100% of simulated 24-observer experiments and by the adjusted model in 5.5%. One route remains partly open: a pure response-history observer leaves a small significant bias in the adjusted stimulus-history amplitude (+0.16° ± 0.06), attributable to the observer conditioning on its own percept rather than the observable target. We report power curves across a range of amplitudes, a specification curve, posterior predictive checks and a failure-mode taxonomy with computable detectors, and release the platform, protocol and preregistration-ready analysis plan. No human data were collected.

1. Introduction

Fischer and Whitney (2014) showed that reports of orientation are biased toward the orientation seen on the previous trial, with a tuning profile that peaks at intermediate differences and vanishes for very different stimuli. The effect has since been reported across many visual dimensions (Kiyonaga et al., 2017; Manassi, Murai & Whitney, 2023), interpreted as a continuity field that stabilises perception (Manassi & Whitney, 2024), and, in a large mega-analysis, found to worsen rather than improve perceptual decisions (Ozkirli, Chetverikov & Pascucci, 2025).

Its origin is contested. Pascucci et al. (2019) found that previous responses predict errors better than previous stimuli, and Gallagher and Benton (2024) characterise the attraction as a pull toward the prior response combined with repulsion from the prior stimulus. Fritsche, Mostert and de Lange (2017), Ceylan, Herzog and Pascucci (2021) and Shan, Hajonides and Myers (2025) locate attractive serial dependence at post-perceptual or decision stages. Fritsche, Spaak and de Lange (2020) show concurrent attraction and repulsion at different timescales, explained by efficient encoding with Bayesian decoding (Wei & Stocker, 2015). Separately, a large literature on magnitude estimation describes regression toward the mean of the stimulus distribution (Jazayeri & Shadlen, 2010; Petzschner, Glasauer & Stephan, 2015), and Wang et al. (2023) show that a single prior-updating mechanism can produce both central tendency and serial dependence. Willemsen et al. (2026) show that stimulus autocorrelation distorts simple-regression estimates of both.

The methodological consequence is that the quantity most papers report — the amplitude of a tuning curve fitted to error against previous-minus-current stimulus — is not, on its own, evidence for perceptual serial dependence. It is an estimator whose false-positive behaviour under the competing accounts has, to our knowledge, not been benchmarked in one place with the design needed to separate them. This paper does that. Its contribution is a platform and a benchmark, not a finding about people.

2. Related work and novelty

research/literature_matrix.csv and research/novelty_memo.md list what is established and what is claimed. Established: the effect, its tuning, the orthogonal-report manipulation, the response-history and decisional accounts, the Bayesian and efficient-coding observers, central tendency, and the autocorrelation confound. Claimed here: a single open benchmark that runs one estimator against a family of history-free generators and reports its false-positive rate; joint identifiability of stimulus- and response-history amplitudes as a function of design; and a browser experiment wired to the same provenance system as the analysis. The novelty claim rests on a targeted, not systematic, search.

3. Experimental paradigm

Fixation (300 ms) → Gabor (500 ms nominal, measured) → blank delay (200 or 1500 ms, blocked) → report cue SAME or ORTHOGONAL → orientation dial from a random start angle → ITI (700 or 1400 ms). Contrast (0.90 or 0.12 with pixel noise) is counterbalanced within cue within block; 12 practice trials with feedback precede six blocks of 100. Orientation is period-180. Stimulus sequences use balanced lag-1 differences so that every similarity level is sampled equally and sequence autocorrelation is zero by construction. Every session is a pure function of an integer seed. Full protocol: research/experiment_protocol.md; data dictionary: docs/DATA_DICTIONARY.md.

The report cue is shown only after stimulus offset, so encoding cannot differ by cue. On cue-incongruent consecutive pairs the response-history predictor is shifted 90° from the stimulus-history predictor; this is what makes the two separable (Section 6.2).

4. Hypotheses

H1 (primary): the group-mean lag-1 stimulus-history amplitude in the adjusted model differs from zero; two-sided, direction not assumed. H2–H4 concern tuning shape, lag decay and sensory uncertainty. H5 and H6 are properties of the design and pipeline — joint identifiability and no false positives from history-free generators — and are the hypotheses this paper tests. Full statements and interpretation rules: research/hypotheses.md, research/preregistration.md.

5. Computational models

All models predict the mean signed circular error with a wrapped-normal likelihood mixed with a uniform lapse. The derivative-of-Gaussian tuning curve is scaled so its amplitude is the peak bias in degrees. The ladder: M0 no history; M0adj nuisance-only; M1 lag-1 DoG on the previous stimulus (the standard analysis); M1lin linear slope; M2 DoG on the previous response; M3 both; M3adj both plus cardinal attraction, running-mean regression at 5 and 20 trials and dial-start pull (primary); M4 lags 1–3 with geometric decay; M5 free uncertainty multiplier; M6 Bayesian observer — circular posterior mean of a von Mises likelihood and a von Mises prior on the previous stimulus, with the likelihood precision tied to the fitted noise so that uncertainty modulation is predicted, not fitted; M7 history mixture. Fitting is per-observer maximum likelihood; group inference is an observer-level BCa bootstrap; model comparison is leave-one-block-out held-out log-likelihood on identical trial sets. Specification: research/model_specification.md.

6. Simulation validation

Synthetic observers (research/simulation_protocol.md) implement the generative process with stimulus history acting on the percept and response history on the motor output. Every study is seeded; results are independent of worker count.

6.1 Parameter recovery

Across a ∈ [−4°, 4°] and sessions of 300, 600, 1200 trials, the largest absolute bias of the adjusted amplitude is 0.18° at 300 / 0.22° at 600 / 0.24° at 1200 trials. Precision is what trial count buys: RMSE at a = 0 is 1.09° at 300 / 0.88° at 600 / 0.60° at 1200 trials. Individual amplitudes from 300-trial sessions should not be interpreted; inference is at the group level.

6.2 Identifiability and design

Jointly recovering a_stim and a_resp on a 3 × 3 grid of true values gives a recovery-error correlation of r = −0.40 (H5 requires |r| < 0.5), with mean bias +0.03° and −0.16°. The design study (60 observers per cell) shows why: with no orthogonal trials the correlation is −0.88 and joint RMSE 1.69°; with half the trials orthogonal, −0.44 and 0.71°. Sequence structure mattered little: balanced differences (0.71°) did not beat i.i.d. uniform sampling (0.70°); they are retained because they make the autocorrelation exactly zero rather than zero in expectation. This is a negative result about a design choice and is reported as one.

sequencep(orthogonal)RMSE a_stimRMSE a_respr(recovery errors)joint RMSE
iid0.500.5130.472−0.430.697
balanced_delta0.500.5160.484−0.440.707
autocorrelated0.500.4950.550−0.450.740
autocorrelated0.250.4710.572−0.490.741
iid0.250.6330.610−0.480.878
balanced_delta0.250.7050.607−0.610.931
autocorrelated0.000.8810.853−0.861.226
iid0.001.0471.129−0.821.540
balanced_delta0.001.1741.219−0.881.692

6.3 Red team: manufacturing serial dependence from nothing

Five generators with a_stim = 0 by construction were each run 200 times on the balanced design and on an autocorrelated random walk; 24-observer experiments were resampled from each pool. The false-positive rate is decomposed into estimator bias (pool mean, t-test) and test calibration (resampling from the re-centred pool).

generator (a_stim = 0 by construction)modelFPR (24 observers)calibrated FPRbias (°) ± SEbias pH6 gate
cardinalM1_stim5.5%5.6%−0.051 ± 0.070= 0.466
cardinalM3_both33.0%5.5%−0.322 ± 0.071< 0.001
cardinalM3adj_full9.2%5.5%−0.097 ± 0.061= 0.114pass
central_tendencyM1_stim100.0%4.8%+1.310 ± 0.043< 0.001
central_tendencyM3_both100.0%5.7%+1.348 ± 0.039< 0.001
central_tendencyM3adj_full5.5%4.8%−0.024 ± 0.062= 0.702pass
motor_persistenceM1_stim5.9%4.5%+0.064 ± 0.088= 0.468
motor_persistenceM3_both4.6%4.6%+0.027 ± 0.092= 0.774
motor_persistenceM3adj_full9.3%4.7%−0.106 ± 0.059= 0.072pass
nullM1_stim9.7%5.1%−0.114 ± 0.059= 0.054
nullM3_both7.0%4.6%−0.065 ± 0.063= 0.304
nullM3adj_full8.8%4.9%−0.103 ± 0.060= 0.086pass
response_historyM1_stim99.9%6.9%+1.086 ± 0.065< 0.001
response_historyM3_both13.6%5.0%+0.158 ± 0.061= 0.010
response_historyM3adj_full13.6%5.6%+0.159 ± 0.061= 0.009FAIL

The standard lag-1 model declares the central-tendency observer serially dependent in 100% of experiments and the response-history observer in 100%; the adjusted model brings the central-tendency rate to 5.5% with bias −0.02° (p = 0.702). The adjusted model does not pass H6 for every generator. It fails for response_history (bias +0.16° ± 0.06, p = 0.009). The response-history leak has an identified cause: the synthetic observer is attracted to its previous response relative to its own noisy percept, whereas the model conditions on the observable report target. This errors-in-variables gap cannot be closed with behavioural data and is carried as a limitation rather than tuned away. On the autocorrelated sequence the null observer alone produces a 5% false-positive rate under the naive model, which is why the pipeline diagnoses sequence autocorrelation before any analysis.

6.4 Model discrimination

Data were generated from eight observers and fitted with eight candidate models. AIC selects the exact generating model in 81% of simulations and a single held-out split in 57%. The coarser and more important verdict — whether the winning model contains a stimulus-history term when, and only when, the generator has one — is correct in 91% (AIC) and 72% (held-out) of simulations. Confusion matrix: results/model_confusion.parquet, figures/fig12_model_confusion.png.

6.5 Power

For a ∈ {0, 0.5, 1, 1.5, 2, 3}° and 300 or 600 trials, 250 fitted observers per cell were resampled into studies of N observers. At a = 0 the rejection rate ranges from 4.3% to 6.0%. With 600 trials, 80% power is reached at N = 24 for 0.5°, 8 for 1.0°, 8 for 1.5°, 8 for 2.0°, 8 for 3.0°. No single effect size is assumed for a human study.

true a (°)trialsN = 8N = 12N = 16N = 20N = 24N = 30N = 40N = 60
0.03006%6%5%5%5%6%5%5%
0.06006%5%4%4%5%5%6%6%
0.530028%38%47%57%64%74%86%95%
0.560040%52%66%74%83%90%96%99%
1.030079%90%95%98%99%100%100%100%
1.060095%100%100%100%100%100%100%100%
1.530097%100%100%100%100%100%100%100%
1.5600100%100%100%100%100%100%100%100%
2.0300100%100%100%100%100%100%100%100%
2.0600100%100%100%100%100%100%100%100%
3.0300100%100%100%100%100%100%100%100%
3.0600100%100%100%100%100%100%100%100%

7. Reference-cohort analysis (SYNTHETIC)

To exercise the full path, a reference cohort of 30 synthetic observers (preset mixed: a_stim mean 2.5° with between-observer SD 1.2°, a_resp 1.2°, running-mean pull 0.05, cardinal 2°, dial pull 0.05, three lags with ρ = 0.4) was analysed exactly as a human dataset would be.

Primary outcome. a_stim = +2.56° (95% CI +2.09 to +2.99; bootstrap p < 0.001; permutation p = 0.010 at resolution 1/101), from 17450 trials. a_resp = +1.38° (+1.25 to +1.51).

Lags. M4 recovers ρ = 0.40 (0.31 to 0.50).

Uncertainty. The cohort has no uncertainty modulation; the paired high − low difference is +0.16° (−0.44 to +0.72, p = 0.587), correctly null.

Model comparison. Best held-out model: M4_multilag, +0.170 nats/trial over M0 for 30/30 observers.

modelΔ held-out LL / trial vs M095% CIobservers improvedparameters
M4_multilag+0.1702+0.1600 to +0.181130 / 3013
M3adj_full+0.1661+0.1563 to +0.175730 / 3012
M5_uncertainty+0.1656+0.1564 to +0.174930 / 3013
M6_bayes+0.1547+0.1467 to +0.163430 / 3011
M3_both+0.0293+0.0229 to +0.036228 / 308
M1_stim+0.0287+0.0223 to +0.035828 / 306
M2_resp+0.0080+0.0061 to +0.010328 / 306
M1lin_stim+0.0045+0.0028 to +0.006622 / 305
M0_no_history+0.0000+0.0000 to +0.00000 / 304

Posterior predictive checks. Generative replication (the response-history predictor recomputed from replicated responses) against three discrepancies:

modelstatisticmean observedmean replicatedmean posterior-predictive p
M0_no_historycurve amplitude2.630.000.03
M0_no_historyerror kurtosis12.9712.650.49
M0_no_historyerror sd10.1510.090.46
M1_stimcurve amplitude2.632.560.48
M1_stimerror kurtosis12.9712.800.51
M1_stimerror sd10.1510.100.47
M3adj_fullcurve amplitude2.632.550.47
M3adj_fullerror kurtosis12.9713.840.59
M3adj_fullerror sd10.1510.580.69
M6_bayescurve amplitude2.631.950.22
M6_bayeserror kurtosis12.9713.860.59
M6_bayeserror sd10.1510.570.69

8. Robustness

Across 45 comparable specifications (nuisance sets × response history × RT cuts × strata × block subsets), 100% share the sign of the primary estimate (+2.56°). Leaving any single observer out moves the estimate by at most 0.10°. Specification curve: figures/fig14_specification_curve.png.

9. Failure analysis

3 of 13 failure detectors fire on the reference cohort — the three that should (motor persistence, central tendency and history confounding are all present in the generator) and none that should not. Four pipeline bugs were found by the validation studies during development and are fixed and tested: unequal trial sets across compared models; running-mean leakage at block starts; a BCa interval excluding its own estimate at small n; and stratified fits that redefined the previous trial. Taxonomy: research/failure_modes.md.

10. Discussion

The estimator most of the field reports is easy to fool. On these simulations a running-mean regression with no trial-level memory produces a significant lag-1 tuning amplitude essentially every time; adjusting for it removes the false positive without removing a real effect. The orthogonal-report cue is not optional: without it the perceptual and response accounts are, in the fitted model, the same parameter. And the adjustment is not complete: a pure response-history process leaves a small bias in the stimulus-history amplitude that no behavioural model can remove, because the observer's reference (its percept) is unobservable. A human amplitude of the size of that bias should not be called perceptual.

Two further cautions. The adjusted model separates a tuned lag-1 component from a linear multi-trial pull; per Wang et al. (2023) these may be one mechanism, and the separation is statistical. And stimulus versus response history is not perception versus decision (Fritsche et al., 2017; Shan et al., 2025).

11. Limitations

See research/limitations.md: no human data; a finite, chosen family of confound generators; the residual response-history leak; statistical rather than mechanistic separation; browser timing and uncalibrated contrast; the mental rotation on orthogonal trials; one perceptual dimension; a non-exhaustive model ladder; discrimination accuracy as the ceiling on model-comparison confidence.

12. Methods summary

Python 3.11, numpy/scipy/pandas (versions in results/release.json); L-BFGS-B with parameter preconditioning and 40 screened starts; observer-level BCa bootstrap with percentile fallback; leave-one-block-out CV; within-block permutation. Browser: Next.js, canvas Gabor precomputed before onset, rAF-timed presentation, measured durations. Exclusion policy exclusion-policy-1.0.0. Reproduce: make reproduce.

13. Ethics

No human participants. No ethics approval sought or claimed. No institutional affiliation claimed. The consent page is labelled as demonstration infrastructure. See research/ethics.md.

14. Conclusion

TimeMind is a ready instrument and a benchmark of its own estimator, with one honest gap. It is not a result about human perception, and it is built so that it cannot be mistaken for one until the study status changes for a real reason.

References

  • Fischer J, Whitney D (2014). Serial dependence in visual perception. Nature Neuroscience 17(5):738-743. https://doi.org/10.1038/nn.3689
  • Cicchini GM, Mikellidou K, Burr D (2017). Serial dependencies act directly on perception. Journal of Vision 17(14):6. https://doi.org/10.1167/17.14.6
  • Fritsche M, Mostert P, de Lange FP (2017). Opposite effects of recent history on perception and decision. Current Biology 27(4):590-595. https://doi.org/10.1016/j.cub.2017.01.006
  • Pascucci D, Mancuso G, Santandrea E, Della Libera C, Plomp G, Chelazzi L (2019). Laws of concatenated perception: Vision goes for novelty, decisions for perseverance. PLOS Biology 17(3):e3000144. https://doi.org/10.1371/journal.pbio.3000144
  • Fritsche M, Spaak E, de Lange FP (2020). A Bayesian and efficient observer model explains concurrent attractive and repulsive history biases in visual perception. eLife 9:e55389. https://doi.org/10.7554/eLife.55389
  • Gallagher GK, Benton CP (2024). Characterizing serial dependence as an attraction to prior response. Journal of Vision 24(9):16. https://doi.org/10.1167/jov.24.9.16
  • Shan J, Hajonides JE, Myers NE (2025). Attractive serial dependence arises during decision-making. PLOS Biology 23(8):e3003333. https://doi.org/10.1371/journal.pbio.3003333
  • Ozkirli A, Chetverikov A, Pascucci D (2025). Large-scale mega-analysis indicates that serial dependence deteriorates perceptual decision-making. Nature Human Behaviour. https://www.nature.com/articles/s41562-025-02362-8
  • Manassi M, Murai Y, Whitney D (2023). Serial dependence in visual perception: A meta-analysis and review. Journal of Vision 23(8):18. https://jov.arvojournals.org/article.aspx?articleid=2791439
  • Manassi M, Whitney D (2024). Continuity fields enhance visual perception through positive serial dependence. Nature Reviews Psychology. https://www.nature.com/articles/s44159-024-00297-x
  • Cicchini GM, Mikellidou K, Burr DC (2024). Serial Dependence in Perception. Annual Review of Psychology. https://doi.org/10.1146/annurev-psych-021523-104939
  • Kiyonaga A, Scimeca JM, Bliss DP, Whitney D (2017). Serial dependence across perception, attention, and memory. Trends in Cognitive Sciences. https://pubmed.ncbi.nlm.nih.gov/28549826/
  • Bliss DP, Sun JJ, D'Esposito M (2017). Serial dependence is absent at the time of perception but increases in visual working memory. Scientific Reports 7:14739. https://doi.org/10.1038/s41598-017-15199-7
  • Ceylan G, Herzog MH, Pascucci D (2021). Serial dependence does not originate from low-level visual processing. Cognition 212:104709. https://doi.org/10.1016/j.cognition.2021.104709
  • Sheehan TC, Serences JT (2022). Attractive serial dependence overcomes repulsive neuronal adaptation. PLOS Biology 20(9):e3001711. https://doi.org/10.1371/journal.pbio.3001711
  • Urai AE, de Gee JW, Tsetsos K, Donner TH (2019). Choice history biases subsequent evidence accumulation. eLife 8:e46331. https://doi.org/10.7554/eLife.46331
  • Jazayeri M, Shadlen MN (2010). Temporal context calibrates interval timing. Nature Neuroscience 13:1020-1026. https://doi.org/10.1038/nn.2590
  • Petzschner FH, Glasauer S, Stephan KE (2015). A Bayesian perspective on magnitude estimation. Trends in Cognitive Sciences 19:285-293. https://doi.org/10.1016/j.tics.2015.03.002
  • Wang T, Luo Y, Ivry RB, Tsay JS, Poppel E, Bao Y (2023). A unitary mechanism underlies adaptation to both local and global environmental statistics in time perception. PLOS Computational Biology 19(5):e1011116. https://doi.org/10.1371/journal.pcbi.1011116
  • Willemsen SCMJ, Oostwoud Wijdenes L, van Beers RJ, Koppen M, Medendorp WP (2026). Does stimulus order affect central tendency and serial dependence in vestibular path integration?. iScience. https://doi.org/10.1016/j.isci.2026.114772
  • Wei XX, Stocker AA (2015). A Bayesian observer model constrained by efficient coding can explain 'anti-Bayesian' percepts. Nature Neuroscience 18:1509-1517. https://www.nature.com/articles/nn.4105
  • Berens P (2009). CircStat: A MATLAB Toolbox for Circular Statistics. Journal of Statistical Software 31(10):1-21. https://doi.org/10.18637/jss.v031.i10