Public Engineering & Reliability Record
Mathematical Validation
We test the math. We publish the results.
Across this 110,000-problem fresh blind validation campaign, Pythos delivered 110,000 verified-correct answers with zero incorrect answers delivered. We publish our complete validation results, methodology, and operational telemetry below.
110,000
Problems Tested
Across 11 STEM domains
110,000
Verified Correct
100.00% verified rate
0
Safely Withheld
0.00% safely withheld
0
Incorrect Delivered
0.00% delivered error rate
🛡️ Current Validation
100.00% of benchmark problems produced a verified correct answer. Across this 110,000-problem fresh blind validation campaign, Pythos delivered 110,000 verified-correct answers with zero incorrect answers delivered. All 11 mathematical domains completed at 100% verified correctness.
| Reliability Metric |
Observed Value |
Percentage / Rate |
Operational Target |
| Problems Tested |
110,000 |
100.00% |
Statistically significant large-scale validation |
| Verified Correct |
110,000 |
100.00% |
High deterministic resolution rate |
| Safely Withheld (UNKNOWN) |
0 |
0.00% |
Safe suppression when confidence is incomplete |
| Incorrect Delivered Answers |
0 |
0.00% |
Zero incorrect answers delivered |
| Verification Escapes |
0 |
0.00% |
Zero unverified mathematical claims escaped |
| False-Positive Rejections |
0 |
0.00% |
Zero valid mathematical proofs rejected |
| Observed Delivered Error Rate |
0.00% (95% Wilson Score CI: [0.0000%, 0.0034%]) |
| Execution Telemetry |
Runtime: 37,646.67s (~10.45 hrs) • Peak Node Heap: 132.1 MB • CAS Concurrency: ≤4 workers steady-state • Provider Failures: 0 • Local LLM Inference: 0% |
Domain Performance Breakdown (110,000 Problems)
| Mathematical Category |
Tested |
Verified Correct |
Safely Withheld |
Incorrect |
Accuracy Status |
| Arithmetic |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Fractions / Decimals / Percentages |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Linear Equations |
10,000 |
10,000 |
0 |
0 |
100.00% Verified (Repaired Router) |
| Systems of Equations |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Quadratics & Polynomials |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Functions & Algebra |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Geometry |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Trigonometry |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Calculus (Derivatives & Integrals) |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Probability & Statistics |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
| Physics / Mechanics |
10,000 |
10,000 |
0 |
0 |
100.00% Verified |
Dedicated Specialized Validation Suites
In addition to the 110,000-problem synthetic mathematical benchmark, all specialized architectural validation suites passed with 100% compliance:
| Validation Suite |
Test Focus |
Cases Passed |
Status |
| Vision Verification |
Geometric diagrams, function graphs, coordinate plots, textbook images |
14 / 14 |
100.00% Passed |
| Safe Withholding UX |
Explicit UNKNOWN messaging, 0 hallucination delivery, pedagogical hints |
9 / 9 |
100.00% Passed |
| Kid Safety & Adversarial |
Self-harm suppression, jailbreak resistance, age-appropriate guardrails |
13 / 13 |
100.00% Passed |
| Multi-Turn Context |
Follow-up questions, pronoun resolution, variable context persistence |
5 / 5 |
100.00% Passed |
| Backup-Brain Recovery |
Automated failover, provider timeouts, 5xx faults, verification invariance |
24 / 24 |
100.00% Passed |
| Project Knowledge & Self-Grounding |
Dynamic website fact retrieval, creator attribution, zero prompt leak |
56 / 56 |
100.00% Passed |
| Accessibility & Math Rendering |
KaTeX math, screen reader ARIA names, focus management, color contrast |
254 / 254 |
100.00% Passed |
| Repaired Linear Router |
Fractions, parentheses, multi-variable, signed terms, edge cases |
33 / 33 |
100.00% Passed |
| End-to-End Verifier |
AST evaluation, SymPy CAS execution, prompt fidelity, delivery gates |
14 / 14 |
100.00% Passed |
🏛️ Historical Validation Evidence
Pythos maintains permanent public records of all completed validation campaigns. Previous frozen benchmarks are preserved to document system progression, architectural hardening, and verification integrity over time.
| Campaign |
Model / Environment |
PRNG Seed |
Tested |
Verified Correct |
Safely Withheld |
Incorrect Delivered |
Verified Rate |
| v1.8.9 Fresh Online Validation |
Groq Cloud (openai/gpt-oss-20b) + Local CAS |
2718281828 |
110,000 |
110,000 |
0 |
0 |
100.00% |
| v1.8.4 Frozen Benchmark |
Deterministic Verifier (Math.js + SymPy) |
1618033988 |
50,000 |
47,907 |
2,093 |
0 |
95.81% |
Environmental / Infrastructure Context (Local Model OOM Event): During an intermediate 110,000-problem benchmark attempt using a locally hosted 7B model (qwen2.5-coder:7b via Ollama on laptop hardware), the local runtime experienced an Out-Of-Memory (OOM) operating system termination (exit code -536870904). This was strictly an environmental and resource failure of running multi-gigabyte local LLM inference under heavy continuous load on constrained laptop hardware—not a mathematical failure of the Pythos verification stack (zero mathematical errors were delivered prior to the OS crash). The campaign was subsequently restructured as a 100% remote online validation campaign using Groq Cloud, completing all 110,000 problems with 132.1 MB peak memory. The crashed run's checkpoints and diagnostic logs are archived in preserved-crashed-local-run/.
Engineering Integrity Notice: While Pythos achieved a 100.00% verified delivery rate across this 110,000-problem blind benchmark with zero incorrect answers delivered, a finite benchmark does not guarantee universal mathematical correctness across every possible future question. Pythos is engineered to systematically verify candidate reasoning and safely withhold answers whenever proof or confidence is incomplete.
🛡️ What does "Safely Withheld" mean?
Sometimes Pythos cannot establish that an answer is sufficiently reliable to present it as fact. In those cases, Pythos may return UNKNOWN rather than guess.
A safely withheld answer is not counted as a correct answer. It is not counted as an incorrect answer either. It means Pythos chose not to deliver an answer that it could not adequately verify.
VERIFIED
The answer passed mathematical verification and was delivered to the student.
SAFELY WITHHELD
Pythos could not establish sufficient confidence, so the answer was withheld to prevent hallucination.
INCORRECT
The delivered answer was mathematically wrong. (Current benchmark result: 0).
The engineering goal is to systematically increase the VERIFIED rate while keeping INCORRECT delivered answers at zero.
⚙️ How Validation Works
Pythos uses an automated verification pipeline that couples generative language models with deterministic Computer Algebra Systems (CAS) and exact rational arithmetic:
1
Generated Mathematical Problem
Seeded pseudorandom generators synthesize unique problems across arithmetic, algebra, calculus, and physics with exact ground-truth solutions.
2
Pythos Receives Problem
The system analyzes deterministic intent, conversational context, and dimensional units to generate candidate step-by-step reasoning.
3
Candidate Mathematical Answer
All quantitative claims, intermediate expressions, and final boxed results are extracted into structured verifiable assertions.
4
Deterministic Multi-Tier Verification
Claims are audited through an in-process exact Math.js AST engine, followed by a bounded SymPy CAS subprocess for deep symbolic calculus and statistics.
5
Strict Prompt-to-Claim Fidelity Gate
The verification bridge inspects whether the mathematical claim answers the actual user prompt rather than a distorted or mutated subexpression.
6
Delivery Decision Gate
If all claims are verified and prompt fidelity is intact, the response is delivered. If any claim is invalid or unprovable, the answer is safely withheld.
Historical Scope & Verification Evolution: The 50,000-problem benchmark documented above is a frozen, authoritative evaluation of the Pythos v1.8.4 deterministic text-verification engine. Subsequent Pythos releases have expanded the verification architecture beyond this milestone—incorporating vision verification for geometric diagrams and textbook photos, alongside verified backup-brain recovery. Newer capabilities build directly on the frozen benchmark foundation without altering historical evidence.
Open Science & Public Evidence
In accordance with our commitment to transparency, the entire benchmark harness, validation scripts, and raw execution telemetry are maintained in our public repository: