Simulation only. No human participants have been recruited or run. Every figure on this site is computed from synthetic observers with known ground truth, and says so.

How many observers, how many trials

Power here is a property of the design and the estimator under an assumed effect size. It is not a promise about a real study.

For each amplitude and session length, 250 independent synthetic observers were simulated and fitted. A study of N observers was then formed by resampling N of them and running the same observer-level test as the primary analysis, 4,000 times. At a true amplitude of zero the rejection rate across all cells ranges from 4.3% to 6.0% (nominal 5%); resampling from a finite pool treats the pool’s own mean as the truth, so these are slightly conservative about calibration.

With 600 trials per observer, 80% power is reached at 24 observers for 0.5°; 8 observers for 1.0°; 8 observers for 1.5°; 8 observers for 2.0°; 8 observers for 3.0°. The literature reports amplitudes over a wide range; the grid spans it rather than betting on one value, and nothing here should be read as the effect size a human study will find.

100%probability of detecting the effect with the correct sign

Fig. 1 Power by number of observers, 300 trials each

SIMULATED DATA

  • 0.5°
  • 1.5°
00.050.20.40.60.81102030405060observersP(detect, correct sign)
Each line is one true amplitude; darker is larger. The grey line at the bottom is the null (a = 0), where the curve is a false-positive rate. Dashed-free rules: 80% and 5%.Source: results/power.parquet · run power-8964a2e-5000
Table view (48 rows)
true a (°)observerspowerany rejection
0.0086.0%6.0%
0.00125.6%5.6%
0.00165.3%5.3%
0.00205.4%5.4%
0.00245.0%5.0%
0.00305.6%5.6%
0.00405.1%5.1%
0.00605.1%5.1%
0.50828.1%28.1%
0.501238.0%38.0%
0.501647.4%47.4%
0.502057.1%57.1%
0.502463.5%63.5%
0.503073.6%73.6%
0.504085.7%85.7%
0.506095.3%95.3%
1.00878.5%78.5%
1.001290.0%90.0%
1.001694.5%94.5%
1.002097.9%97.9%
1.002499.0%99.0%
1.003099.9%99.9%
1.0040100.0%100.0%
1.0060100.0%100.0%
1.50897.2%97.2%
1.501299.8%99.8%
1.5016100.0%100.0%
1.5020100.0%100.0%
1.5024100.0%100.0%
1.5030100.0%100.0%
1.5040100.0%100.0%
1.5060100.0%100.0%
2.00899.7%99.7%
2.0012100.0%100.0%
2.0016100.0%100.0%
2.0020100.0%100.0%
2.0024100.0%100.0%
2.0030100.0%100.0%
2.0040100.0%100.0%
2.0060100.0%100.0%
3.008100.0%100.0%
3.0012100.0%100.0%
3.0016100.0%100.0%
3.0020100.0%100.0%
3.0024100.0%100.0%
3.0030100.0%100.0%
3.0040100.0%100.0%
3.0060100.0%100.0%

Fig. 2 Power by number of observers, 600 trials each

SIMULATED DATA

  • 0.5°
  • 1.5°
00.050.20.40.60.81102030405060observersP(detect, correct sign)
Each line is one true amplitude; darker is larger. The grey line at the bottom is the null (a = 0), where the curve is a false-positive rate. Dashed-free rules: 80% and 5%.Source: results/power.parquet · run power-8964a2e-5000
Table view (48 rows)
true a (°)observerspowerany rejection
0.0085.5%5.5%
0.00125.1%5.1%
0.00164.5%4.5%
0.00204.3%4.3%
0.00245.0%5.0%
0.00304.7%4.7%
0.00405.9%5.9%
0.00605.9%5.9%
0.50840.0%40.1%
0.501251.6%51.6%
0.501665.7%65.7%
0.502073.9%73.9%
0.502482.5%82.5%
0.503089.6%89.6%
0.504095.9%95.9%
0.506099.4%99.4%
1.00894.5%94.5%
1.001299.6%99.6%
1.0016100.0%100.0%
1.0020100.0%100.0%
1.0024100.0%100.0%
1.0030100.0%100.0%
1.0040100.0%100.0%
1.0060100.0%100.0%
1.508100.0%100.0%
1.5012100.0%100.0%
1.5016100.0%100.0%
1.5020100.0%100.0%
1.5024100.0%100.0%
1.5030100.0%100.0%
1.5040100.0%100.0%
1.5060100.0%100.0%
2.008100.0%100.0%
2.0012100.0%100.0%
2.0016100.0%100.0%
2.0020100.0%100.0%
2.0024100.0%100.0%
2.0030100.0%100.0%
2.0040100.0%100.0%
2.0060100.0%100.0%
3.008100.0%100.0%
3.0012100.0%100.0%
3.0016100.0%100.0%
3.0020100.0%100.0%
3.0024100.0%100.0%
3.0030100.0%100.0%
3.0040100.0%100.0%
3.0060100.0%100.0%