Public Engineering & Reliability Record

Mathematical Validation

We test the math. We publish the results.

Across this 110,000-problem fresh blind validation campaign, Pythos delivered 110,000 verified-correct answers with zero incorrect answers delivered. We publish our complete validation results, methodology, and operational telemetry below.

110,000
Problems Tested
Across 11 STEM domains
110,000
Verified Correct
100.00% verified rate
0
Safely Withheld
0.00% safely withheld
0
Incorrect Delivered
0.00% delivered error rate

🛡️ Current Validation

Engine Version: Pythos v1.8.9
Validation Campaign: Fresh Online 110K (Remote Groq Cloud)
PRNG Stream Seed: 2718281828
Last Validated: September 24, 2026
100.00% of benchmark problems produced a verified correct answer. Across this 110,000-problem fresh blind validation campaign, Pythos delivered 110,000 verified-correct answers with zero incorrect answers delivered. All 11 mathematical domains completed at 100% verified correctness.
Reliability Metric Observed Value Percentage / Rate Operational Target
Problems Tested 110,000 100.00% Statistically significant large-scale validation
Verified Correct 110,000 100.00% High deterministic resolution rate
Safely Withheld (UNKNOWN) 0 0.00% Safe suppression when confidence is incomplete
Incorrect Delivered Answers 0 0.00% Zero incorrect answers delivered
Verification Escapes 0 0.00% Zero unverified mathematical claims escaped
False-Positive Rejections 0 0.00% Zero valid mathematical proofs rejected
Observed Delivered Error Rate 0.00%   (95% Wilson Score CI: [0.0000%, 0.0034%])
Execution Telemetry Runtime: 37,646.67s (~10.45 hrs) • Peak Node Heap: 132.1 MB • CAS Concurrency: ≤4 workers steady-state • Provider Failures: 0 • Local LLM Inference: 0%

Domain Performance Breakdown (110,000 Problems)

Mathematical Category Tested Verified Correct Safely Withheld Incorrect Accuracy Status
Arithmetic 10,000 10,000 0 0 100.00% Verified
Fractions / Decimals / Percentages 10,000 10,000 0 0 100.00% Verified
Linear Equations 10,000 10,000 0 0 100.00% Verified (Repaired Router)
Systems of Equations 10,000 10,000 0 0 100.00% Verified
Quadratics & Polynomials 10,000 10,000 0 0 100.00% Verified
Functions & Algebra 10,000 10,000 0 0 100.00% Verified
Geometry 10,000 10,000 0 0 100.00% Verified
Trigonometry 10,000 10,000 0 0 100.00% Verified
Calculus (Derivatives & Integrals) 10,000 10,000 0 0 100.00% Verified
Probability & Statistics 10,000 10,000 0 0 100.00% Verified
Physics / Mechanics 10,000 10,000 0 0 100.00% Verified

Dedicated Specialized Validation Suites

In addition to the 110,000-problem synthetic mathematical benchmark, all specialized architectural validation suites passed with 100% compliance:

Validation Suite Test Focus Cases Passed Status
Vision Verification Geometric diagrams, function graphs, coordinate plots, textbook images 14 / 14 100.00% Passed
Safe Withholding UX Explicit UNKNOWN messaging, 0 hallucination delivery, pedagogical hints 9 / 9 100.00% Passed
Kid Safety & Adversarial Self-harm suppression, jailbreak resistance, age-appropriate guardrails 13 / 13 100.00% Passed
Multi-Turn Context Follow-up questions, pronoun resolution, variable context persistence 5 / 5 100.00% Passed
Backup-Brain Recovery Automated failover, provider timeouts, 5xx faults, verification invariance 24 / 24 100.00% Passed
Project Knowledge & Self-Grounding Dynamic website fact retrieval, creator attribution, zero prompt leak 56 / 56 100.00% Passed
Accessibility & Math Rendering KaTeX math, screen reader ARIA names, focus management, color contrast 254 / 254 100.00% Passed
Repaired Linear Router Fractions, parentheses, multi-variable, signed terms, edge cases 33 / 33 100.00% Passed
End-to-End Verifier AST evaluation, SymPy CAS execution, prompt fidelity, delivery gates 14 / 14 100.00% Passed

🏛️ Historical Validation Evidence

Pythos maintains permanent public records of all completed validation campaigns. Previous frozen benchmarks are preserved to document system progression, architectural hardening, and verification integrity over time.

Campaign Model / Environment PRNG Seed Tested Verified Correct Safely Withheld Incorrect Delivered Verified Rate
v1.8.9 Fresh Online Validation Groq Cloud (openai/gpt-oss-20b) + Local CAS 2718281828 110,000 110,000 0 0 100.00%
v1.8.4 Frozen Benchmark Deterministic Verifier (Math.js + SymPy) 1618033988 50,000 47,907 2,093 0 95.81%
Environmental / Infrastructure Context (Local Model OOM Event): During an intermediate 110,000-problem benchmark attempt using a locally hosted 7B model (qwen2.5-coder:7b via Ollama on laptop hardware), the local runtime experienced an Out-Of-Memory (OOM) operating system termination (exit code -536870904). This was strictly an environmental and resource failure of running multi-gigabyte local LLM inference under heavy continuous load on constrained laptop hardware—not a mathematical failure of the Pythos verification stack (zero mathematical errors were delivered prior to the OS crash). The campaign was subsequently restructured as a 100% remote online validation campaign using Groq Cloud, completing all 110,000 problems with 132.1 MB peak memory. The crashed run's checkpoints and diagnostic logs are archived in preserved-crashed-local-run/.
Engineering Integrity Notice: While Pythos achieved a 100.00% verified delivery rate across this 110,000-problem blind benchmark with zero incorrect answers delivered, a finite benchmark does not guarantee universal mathematical correctness across every possible future question. Pythos is engineered to systematically verify candidate reasoning and safely withhold answers whenever proof or confidence is incomplete.

🛡️ What does "Safely Withheld" mean?

Sometimes Pythos cannot establish that an answer is sufficiently reliable to present it as fact. In those cases, Pythos may return UNKNOWN rather than guess.

A safely withheld answer is not counted as a correct answer. It is not counted as an incorrect answer either. It means Pythos chose not to deliver an answer that it could not adequately verify.

VERIFIED
The answer passed mathematical verification and was delivered to the student.
SAFELY WITHHELD
Pythos could not establish sufficient confidence, so the answer was withheld to prevent hallucination.
INCORRECT
The delivered answer was mathematically wrong. (Current benchmark result: 0).

The engineering goal is to systematically increase the VERIFIED rate while keeping INCORRECT delivered answers at zero.

⚙️ How Validation Works

Pythos uses an automated verification pipeline that couples generative language models with deterministic Computer Algebra Systems (CAS) and exact rational arithmetic:

1

Generated Mathematical Problem

Seeded pseudorandom generators synthesize unique problems across arithmetic, algebra, calculus, and physics with exact ground-truth solutions.

2

Pythos Receives Problem

The system analyzes deterministic intent, conversational context, and dimensional units to generate candidate step-by-step reasoning.

3

Candidate Mathematical Answer

All quantitative claims, intermediate expressions, and final boxed results are extracted into structured verifiable assertions.

4

Deterministic Multi-Tier Verification

Claims are audited through an in-process exact Math.js AST engine, followed by a bounded SymPy CAS subprocess for deep symbolic calculus and statistics.

5

Strict Prompt-to-Claim Fidelity Gate

The verification bridge inspects whether the mathematical claim answers the actual user prompt rather than a distorted or mutated subexpression.

6

Delivery Decision Gate

If all claims are verified and prompt fidelity is intact, the response is delivered. If any claim is invalid or unprovable, the answer is safely withheld.

Historical Scope & Verification Evolution: The 50,000-problem benchmark documented above is a frozen, authoritative evaluation of the Pythos v1.8.4 deterministic text-verification engine. Subsequent Pythos releases have expanded the verification architecture beyond this milestone—incorporating vision verification for geometric diagrams and textbook photos, alongside verified backup-brain recovery. Newer capabilities build directly on the frozen benchmark foundation without altering historical evidence.

Open Science & Public Evidence

In accordance with our commitment to transparency, the entire benchmark harness, validation scripts, and raw execution telemetry are maintained in our public repository:

Inspect Pythos-Tests Repository ↗