Tests
Add Test
Audience Config
Saved Panels
History
Compare Models
No test selected
Add your known A/B tests, then run the synthetic audience against them to validate calibration.
Add A/B Test
Variant A
Variant B
Ξ < 0.3% β weight 0.5
Ξ 0.3β1% β weight 1.0
Ξ > 1% β weight 1.5β2.0
Ξ 0.3β1% β weight 1.0
Ξ > 1% β weight 1.5β2.0
Add up to 20 known A/B tests. Run all at once to validate calibration.
Audience Config
Autoresearch edits these parameters to improve calibration. Version: β
Scoring
Spam Penalty Weight
Use Median (not mean)
Margin Threshold (% pts)
Prompt Calibration
Default Skepticism
Relevance Sensitivity
Fear Response Level
Trust Prior
Engagement Mix (must sum to 1.0)
Passive
Reactive
Health-Anxious
Active Manager
Age Mix (must sum to 1.0)
24β35
36β50
51β65
66β80
autoresearch.md β Agent Working Memory
Loading...
Saved Audience Panels
Frozen 250-persona instances for deterministic replay scoring
Score New Variant Against Panel
Model Comparison
Accuracy trajectories, plateau points, and cost across all model tracks
Weighted Accuracy β All Models
Best Config Achieved Per Model
| Model | Best Weighted Acc | Best Raw Acc | Iterations | $/Run-All (est) | Wall Clock/Run-All (est) |
|---|
Test Agreement Analysis β which tests do models disagree on?
Run multiple models to see agreement analysis.
Autoresearch Dashboard β β
Every Run All iteration β config changes, accuracy trajectory, test flips
Accuracy Trajectory β Raw / Weighted / Qualified
No iterations yet
Click Run All to start the autoresearch loop.