Calibration Dashboard
Add A/B tests to begin validation
Tests
Add Test
Audience Config
Saved Panels
History
Compare Models
β€”
Raw Accuracy
β€”
Weighted Accuracy
β€”
Qualified Accuracy
β€”
Strong Consensus
β€”
Split Decisions
πŸ“Š

No test selected

Add your known A/B tests, then run the synthetic audience against them to validate calibration.

Add A/B Test
Variant A
Variant B
Ξ” < 0.3% β†’ weight 0.5
Ξ” 0.3–1% β†’ weight 1.0
Ξ” > 1% β†’ weight 1.5–2.0
Add up to 20 known A/B tests. Run all at once to validate calibration.
Audience Config
Autoresearch edits these parameters to improve calibration. Version: β€”
Scoring
Spam Penalty Weight
Use Median (not mean)
Margin Threshold (% pts)
Prompt Calibration
Default Skepticism
Relevance Sensitivity
Fear Response Level
Trust Prior
Engagement Mix (must sum to 1.0)
Passive
Reactive
Health-Anxious
Active Manager
Age Mix (must sum to 1.0)
24–35
36–50
51–65
66–80
autoresearch.md β€” Agent Working Memory
Loading...
Saved Audience Panels
Frozen 250-persona instances for deterministic replay scoring
Score New Variant Against Panel
Model Comparison
Accuracy trajectories, plateau points, and cost across all model tracks
Weighted Accuracy β€” All Models
Best Config Achieved Per Model
Model Best Weighted Acc Best Raw Acc Iterations $/Run-All (est) Wall Clock/Run-All (est)
Test Agreement Analysis β€” which tests do models disagree on?
Run multiple models to see agreement analysis.
Autoresearch Dashboard β€” β€”
Every Run All iteration β€” config changes, accuracy trajectory, test flips
Accuracy Trajectory β€” Raw / Weighted / Qualified
πŸ“‹

No iterations yet

Click Run All to start the autoresearch loop.