Writing
The Benchmark–Intelligence Gap: Why High Scores Overstate Fluid and Commonsense Intelligence
Benchmark scores measure displayed skill on fixed tasks, not efficient learning under novelty. A position paper on ARC-AGI, commonsense evaluation, and what honest measurement would require.
Frontier Language Models in September 2026: Truthfulness, Sycophancy, and Code Quality
Snapshot: September 22, 2026. Gemini 3.8 Flash, Claude Sonnet 5, Muse Spark 1.3, and DeepSeek V4 Pro 0813 measured against truthfulness, presuppositional integrity, sycophancy, and forensic code quality. No model earns trust.