pruuf.
Admin
Checking your access…
⚠
Applies to every article PRUUF analyses. The PRUUF judge scores every article readers see; these settings choose the models around it — research, the writer, and the teachers that label its training data. The provider and model for each phase must be compatible — the model dropdown only lists models valid for the selected provider. Make sure the required API key is set in config/secrets.
Production PRUUF Judge
Scores every article readers see: the eight categories, the opinion share and the summary, from the article and the verified research. Llama 3.1 8B with PRUUF's own fine-tuned adapter, served from PRUUF's GPU on Cloud Run. The adapter changes only when model/promote_judge.py finds a retrained one that beats it on the frozen holdout.
…
…
…
Phase 1 Research
Verifies the article's factual claims against the live web before scoring. Anthropic uses the model's built-in web search (Claude only). Tavily fetches web snippets, then the chosen model synthesizes the verdicts.
Writer Takeaway, topic and category notes
The judge scores and summarises but was never trained to write the panel's takeaway, to pick the topic that shows the health, finance or legal note, or to explain a category when a reader opens it. This model writes those from the judge's result. It never changes a score.
Teachers Lead
Behind the scenes only. After the judge has scored an article and the reader has their result, the teacher panel scores the same article again to make a training label for the next judge. The lead writes that label's summary. Nothing the teachers produce reaches a reader or changes a score.
Teachers Panel
Each member scores the article independently and the per-category median becomes the training label for PRUUF's judge — never a score a reader sees. Three members is the minimum that does real work: a median of two is just their midpoint, so a dissenting member is averaged in at full weight instead of being outvoted. Every member added multiplies labelling cost, and switching the panel off stops new training labels without affecting readers at all.
Nightly Batch overrides
The overnight news agent uses the settings above by default; the judge scores its articles too. Enable this to give the nightly batch its own research model and teacher lead (for example a cheaper one) without changing interactive analysis.
Reading levels Shown on the extension’s profile screen

One row per level, lowest first. At is the lifetime number of checks that earns it. The first row must be 0 — a reader with no checks still needs a level — and thresholds must increase.

Loading…