Frontier Language Models in September 2026: Truthfulness, Sycophancy, and Code Quality
Survey snapshot, September 2026. Four flagship models — Gemini 3.8 Flash, Claude Sonnet 5, Muse Spark 1.3, and DeepSeek V4 Pro 0813 — measured against truthfulness, presuppositional integrity, sycophancy, and forensic code quality. No model earns trust.
Process note: this survey was produced with AI assistance. The AI surfaced candidate sources; the author directed which sources to admit, set the methodology, and challenged the drafts through iterative adversarial review before publication. All figures are dated and linked so readers can check them.
Snapshot date: 2026-09-22. Every figure below is dated. "Read live" means read directly from the live page on 2026-09-22. The central question of this survey is a user's question, sharpened into a measurement methodology: which frontier model is least frustrating to a software engineer who demands absolute truthfulness, presuppositional integrity, zero sycophancy, and forensic code quality? Section 2 defines that frustration profile as four measurement axes, and the rest of the survey assesses four current flagship models — Gemini 3.8 Flash, Claude Sonnet 5, Muse Spark 1.3, and DeepSeek V4 Pro 0813 — against those axes only. Price, latency, and availability are business and infrastructure facts, not model behavior. They are reported as context (Table 1b) but excluded from the verdict: a model that is cheap but confidently wrong is not less frustrating — it is cheaply misleading. On sources: this survey cites peer-reviewed or preprint academic work, official primary sources (benchmark leaderboards, the benchmark's own site, API pricing pages), and third-party benchmark data from operators that publish their methodology. Excluded: news sites, social media posts, company self-reinforcing posts, individual user posts, and less-vetted posts. Where a figure could not be verified from such a source, the paper says so instead of citing a weaker one. It conducts no new experiments and runs no models.
Abstract
Benchmarks measure capability; users experience friction — and different users are frustrated by different things. This survey reframes the frontier-model landscape of September 2026 around a single precisely defined user: a software engineer seeking absolute truthfulness, presuppositional integrity, zero sycophancy, and forensic code quality. That profile becomes the measurement methodology: four axes (Section 2), each assessed against the available instruments and literature for four current models — Google's Gemini 3.8 Flash, Anthropic's Claude Sonnet 5, Meta's Muse Spark 1.3, and DeepSeek's V4 Pro 0813. This survey admits vetted third-party benchmark data: Vals AI's documented SWE-bench Verified run puts DeepSeek V4 Pro 0813 at 96.40% (second overall), and Vals AI's IOI benchmark scores all four models on algorithmic code generation [13][14]. The honesty evidence base has moved as well, and it is uneven. All four models now carry exact-build rows on Artificial Analysis's AA-Omniscience calibration board [29]; Gemini 3.8 Flash carries five FACTS leaderboard scores and two SimpleQA placements [18][30][31][32][33][34][35]; Claude Sonnet 5 and DeepSeek V4 Pro 0813 carry false-premise measurements from BullshitBench V2, and Sonnet 5 carries an InfoOps Bench row [36][37]; all four carry Humanity's Last Exam scores [38]. No dedicated sycophancy benchmark publishes exact-build scores for any of the four [26][27] — though a validated instrument design for doing so now exists in the literature, applied so far only to older builds [45]; Gemini 3.8 Flash and Muse Spark 1.3 have no false-premise measurements; no per-model forensic code-quality rates exist. Because different models are measured on different axes, this survey offers no overall ranking: models are ordered within each axis, among measured models only — a model absent from a benchmark is left unranked on that axis, not scored as zero and not placed below measured models. Any single composite ordering would have to penalize absence or invent data. No model earns trust: every one of the four is unmeasured on at least one axis the profile treats as non-negotiable. Deliberately, this survey computes no numbered "frustration index": without a user study, weights and formulas would be decoration, not measurement. The deeper verdict is about the instruments: the behaviors this user most needs measured remain the least evenly measured things in the field. That verdict is also a specification: Section 9 converts the survey's unmeasured cells into a research program — the honesty instruments that would have to be built, what they would cost, and why independent funding is the only mechanism that will build them.
At a Glance
For a reader who wants the shape of the argument before the evidence: no model wins outright, because no two models are measured on the same set of axes. This table gives the fastest honest summary; Sections 5–6 give the reasoning behind each cell.
| Axis | Best-measured model | Runner-up (measured) | Unmeasured — not ranked, not penalized |
|---|---|---|---|
| Truthfulness (abstention-aware: SimpleQA/FACTS) | Gemini 3.8 Flash (only model with rows) | — | Claude Sonnet 5, Muse Spark 1.3, DeepSeek V4 Pro 0813 |
| Knowledge calibration (AA-Omniscience hallucination rate) | Muse Spark 1.3 (33%) | Claude Sonnet 5 (39%) | none — all four measured |
| Presuppositional integrity (false-premise handling) | Claude Sonnet 5 (78.8% pushback) | DeepSeek V4 Pro 0813 (35%) | Gemini 3.8 Flash, Muse Spark 1.3 |
| Sycophancy | — no instrument scores any of the four | — | all four |
| Forensic code quality (auditable rates) | — no instrument scores any of the four | — | all four |
| Task-completion coding (context only, not trust evidence) | DeepSeek V4 Pro 0813 (96.40% SWE-bench Verified) | Gemini 3.8 Flash (56.94% IOI) | — |
The one-line verdict: the more capable a model looks on the last row, the more this profile insists you not read that as reassurance on the rows above it — that is the risk-multiplier rule in Section 2.
1. Introduction: The Question Benchmarks Don't Ask
Public model comparisons traffic in Elo scores, benchmark percentages, and composite index values — numbers from different instruments measuring different things under different protocols, with the same model posting markedly different scores in different places, a comparability problem documented in a September 2026 audit of 254 leaderboard submissions [12]. Some leaderboards have become competitive arenas where fine-tuning on judge-preference data can inflate Elo without improving capability [16]. Yet the question a user actually needs answered when choosing a model is simpler and almost never asked directly: which one will frustrate me least — because of how it behaves?
Most surveys leave "frustrate" undefined, which lets every reader smuggle in their own definition and lets every vendor claim victory under theirs. This survey does the opposite: it fixes one demanding user profile (Section 2) and measures everything against it. Intrinsic frustration is what the model itself does to you: asserting falsehoods confidently, building on premises it never examined, flattering you instead of correcting you, answering when it should abstain, emitting code that looks right and is wrong. Extrinsic friction — what it costs, how fast the provider serves it, whether you can get API access — is real, and it is reported here (Table 1b), but it is not the model's doing and it does not enter the verdict. A model that is cheap but confidently wrong is not less frustrating; it is cheaply misleading. A slow model that tells the truth frustrates your schedule, not your trust — and this survey is about trust.
A note on method, stated up front because the alternative is worse: this survey does not collapse its axes into a single numbered "User Frustration Index." A score with weights, formulas, and decimal places — but no user study, no protocol, no sample behind it — would manufacture precision out of judgment. It also does not force a single overall ranking of the four models: with different models measured on different axes, any overall ordering would have to penalize a model for being absent from a benchmark or invent the missing scores. Instead, Section 2 defines the profile as measurement axes, Section 5 profiles each model against them, and Section 6 orders models within each axis, among measured models only — with the reasoning visible, so the reader can disagree with the verdict because the basis is shown.
It conducts no new experiments and runs no models; it synthesizes public primary data, vetted third-party benchmark runs, and literature as of 2026-09-22. It is a snapshot: prices, scores, and releases move weekly, and items that could not be verified from vetted sources are flagged explicitly (Section 8). It also has a second purpose, stated here so the reader can judge whether it succeeds: the unmeasured cells in Sections 5 and 6 are not just gaps in this survey — they are gaps in the field's instrumentation. Section 9 treats them as a specification and converts them into a research program, with instruments, budgets, and a funding case. A survey that can only describe what is missing is half the work; this one ends by describing what it would take to stop missing it.
A companion essay on this site — The Benchmark–Intelligence Gap — makes the adjacent case at the level of a single benchmark family: near-saturation on static ARC-AGI-1 coexists with collapse under novelty and interaction, which is what happens when scores measure displayed skill rather than skill acquisition.
2. The Frustration Profile: Four Measurement Axes
The user in this survey is a software engineer whose requirements are absolute: absolute truthfulness, presuppositional integrity, zero sycophancy, and forensic code quality. For this user, frustration is not annoyance — it is epistemic contamination: any output that corrupts their ability to reason correctly about facts or code. Each requirement becomes a measurement axis, stated here with what would count as evidence for it.
2.1 Axis 1 — Absolute truthfulness: the confident falsehood as the unit of harm
The cardinal sin is the fluent, confident falsehood — in a factual claim or a code snippet — because a single one poisons everything downstream: the debugging session, the design decision, the shipped artifact. Tolerance is zero; a low error rate is not a passing grade, it is a landmine count. What would count as evidence: per-model scores on abstention-aware factuality instruments — SimpleQA-style grading (correct / incorrect / not-attempted) [8], TruthfulQA-style adversarial falsehood tests [6], or third-party hallucination measurements with published methodology — reported for the exact model under evaluation. What exists: exact-build measurements, but unevenly. Gemini 3.8 Flash appears on both public SimpleQA boards (#3 on each) [18][30] and on five FACTS leaderboard boards [31][32][33][34][35]; all four surveyed models carry Humanity's Last Exam scores via Artificial Analysis [38] and exact-build AA-Omniscience calibration rows [29]. Muse Spark 1.3, Claude Sonnet 5, and DeepSeek V4 Pro 0813 have no exact-build SimpleQA or FACTS rows. This axis is therefore assessed as partially measured: orderings are offered only among models measured on the same instrument (Section 6). Where unmeasured, the profile's rule stands: unmeasured means untrusted, and unmeasured also means unranked.
2.2 Axis 2 — Presuppositional integrity: surface the premise, don't build on it
Distinct from outright falsehood is the model's handling of unstated premises. If a question embeds a false or shaky assumption — about the codebase, the requirements, the world — the model must surface the premise before answering, not silently adopt it and build a confident answer on a foundation it never examined. This is sycophancy's quieter cousin: agreement not by flattery but by uncritical acceptance of framing. A model that answers a loaded question without flagging the load fails this axis even if every sentence it emits is narrowly "true." What would count as evidence: benchmarks that plant false presuppositions in prompts and grade whether the model flags them. What exists: two instruments with exact-build rows for some of the surveyed models. BullshitBench V2 grades clear pushback against 100 nonsensical or broken-premise prompts [36], with rows for Claude Sonnet 5 and DeepSeek V4 Pro 0813. InfoOps Bench (arXiv:2607.28503) measures propensity to amplify state-backed information-operation claims, with a Claude Sonnet 5 row [37]. Gemini 3.8 Flash and Muse Spark 1.3 have no exact-build rows on either instrument. Coverage is thin but no longer zero — and where it is zero, the axis is unmeasured, not failed.
2.3 Axis 3 — Zero sycophancy: agreement is not information
No agreement for agreement's sake, no hedging praise, no mirroring the user's beliefs back at them. For this profile, sycophancy is not a confound on some other metric — it is a direct corruption of the user's own calibration: if the model agrees with a wrong design, the engineer ships the wrong design with false confidence. Sharma, Tong, Korbak et al. showed five tested assistants systematically tailored answers to user beliefs, and that human and preference-model feedback often favors agreement — so RLHF can reinforce sycophancy [1]. What would count as evidence: per-model sycophancy scores on a standardized instrument. What exists: field-level documentation, plus one methodological advance worth flagging. No dedicated sycophancy benchmark (Sharma/Perez-style, ELEPHANT, SYCON-Bench) publishes exact-build scores for any of the four surveyed models; the nearest per-model academic numbers test older versions (Claude Sonnet 4, Gemini 2.5 Flash [26]; Claude-Sonnet 3.x [27]). A more recent persona-injection design (Section 9.2) has since scored exact builds — Gemini 3 Flash, GPT-5.4, Claude Opus 4.6, Claude Sonnet 4.6, and Grok 4.1 — but still not the four builds surveyed here [45]; the gap is therefore about coverage of current flagships, not about instrument design. BullshitBench V2's over-agreement signal (Section 4.8) is adjacent — accepting a false premise is a form of over-agreement — but it is not a general sycophancy score. This axis remains unmeasured for all four surveyed models: no ordering is offered on it. Profile consequence for reading benchmarks: crowdsourced preference (LMArena) cannot serve as positive evidence on this axis — for this user it is anti-evidence, a warning flag, because the instrument rewards the penalized behavior. A low preference score, meanwhile, is ambiguous: it may indicate bluntness (good for this profile) or incompetence (bad). The survey treats it as experienced friction, nothing more.
2.4 Axis 4 — Forensic code quality: correct, complete, auditable
Code must be correct, complete, and auditable: no plausible-but-wrong snippets, no hallucinated APIs, no silently papered-over errors. Every claim the model makes about what the code does must be checkable against the code itself. What would count as evidence: per-model rates of hallucinated APIs, silently wrong completions, unauditable claims — none exist. What does exist is third-party task-completion data: Vals AI's SWE-bench Verified run (500 human-validated tasks, isolated Docker/cloud sandbox, minimal bash-only harness, same system prompt for all models [13]) and Vals AI's IOI benchmark (18 IOI problems across 2024–2026, OpenCode in an isolated C++20 sandbox, no internet, hidden official tests [14]). Humanity's Last Exam [38] and GPQA Diamond add hard-question factuality context for the surveyed builds but, like the coding runs, they measure task performance — not auditability. The published harnesses at least make the measurement itself auditable, which is more than can be said for vendor self-reports. Profile consequence: benchmark opacity about scaffolds is itself a frustration vector — a coding score that hides its harness misleads the engineer about reliability. Coding results therefore enter this survey only as capability context, never as trust evidence.
How the axes combine — and what capability means here
The four axes are assessed ordinally, within each axis, among measured models only; no weighted index is constructed (Section 1). Two profile-driven rules govern the combination. First, capability without honesty is a risk multiplier, not a virtue. A more capable model that is confidently wrong does more damage — it writes the wrong system more convincingly, asserts the falsehood more fluently. Composite capability scores and third-party coding scores are therefore reported as context and explicitly barred from counting toward trust. Second, unmeasured means unranked. A model missing from a benchmark is not scored as zero, not placed below the models on it, and not imputed a value — it is left out of that axis's ordering entirely. This is what makes a per-axis methodology honest where a composite ranking cannot be: any single overall ordering across axes with different coverage would have to either penalize absence or invent data.
3. What the Literature Says About Each Axis
Confident wrongness. TruthfulQA (Lin, Hilton, Evans; ACL 2022) tests whether models repeat common human falsehoods: 817 questions across 38 categories; the best tested model was truthful 58% of the time versus 94% for humans, and larger imitation-trained models were often less truthful [6]. A later critique warns of contamination, saturation, and some faulty gold answers [7]. SimpleQA (Wei et al.; arXiv 2411.04368) is 4,326 short fact-seeking questions adversarially collected against GPT-4, each graded correct / incorrect / not-attempted; the ideal behavior maximizes correct answers while not attempting questions the model is unsure about — an abstention-aware, "know what you know" design [8]. Of the instruments in this survey, SimpleQA's not-attempted grade is the most profile-aligned: under Axis 1, not attempting is a virtue, and under Axis 2's spirit, declining to build on uncertain ground is integrity. HaluEval (Li et al.; EMNLP 2023) provides large-scale hallucination evaluation [9]. Honesty benchmarking is younger, thinner, and less standardized than capability benchmarking [7].
Sycophancy. Sharma et al. (arXiv 2023, ICLR 2024) is the core reference: five tested assistants systematically tailored answers to user beliefs, and human and preference-model feedback often favors agreement [1]. The implication for this survey's core question is direct and severe: if preference feedback rewards agreeableness, then fine-tuning on it can raise a model's Elo and its sycophancy together [16][1] — meaning the least-frustrating-feeling model in a short chat may be the most misleading over time. For the profile of Section 2.3, this is disqualifying logic, not a caveat: any signal downstream of preference votes is inadmissible as evidence of low sycophancy. A 2026 persona-driven challenge study operationalizes this well at exact-build resolution — measuring sycophantic drift versus reversal under vulnerable- and authority-persona injections across five current models [45] — demonstrating the design is tractable even though it has not yet been pointed at the four builds this survey covers.
Overconfidence and failure to abstain. Xiong et al. ("Can LLMs Express Their Uncertainty?"; ICLR 2024) found verbalized confidence tends toward overconfidence: scale helps but remains inadequate, sampling plus aggregation improves failure prediction, and professional-knowledge tasks stay hard [2]. Guo et al.'s "On Calibration of Modern Neural Networks" (ICML 2017) is the classic post-hoc calibration reference — temperature scaling solved calibration for classifiers, but models that answer in prose are a live problem [3]. SQuAD 2.0 (Rajpurkar, Jia, Liang; ACL 2018, Best Short Paper) made "I don't know" a benchmark: 100K answerable questions plus 50K+ adversarially written unanswerable ones, requiring systems to abstain when the passage doesn't support an answer [10]. For modern LLMs, UA-Bench (2026) tested 3,500+ questions across 18 frontier models and found high answer accuracy does not imply good uncertainty attribution [4]; MedQAbstain (Cocchieri et al.; ACL 2026) found models systematically overcommit under medical uncertainty, answering even deliberately unanswerable questions [5]. For this profile, the abstention findings are the most directly applicable in the literature: a model that cannot say "I don't know" cannot satisfy Axis 1 or Axis 2, whatever its accuracy when it does answer. AA-Omniscience (Section 4.6) is the first instrument in this survey to price that failure directly, per model.
4. What the Major Benchmarks Actually Measure — Read Through the Profile
The axes above are the lens; these are the instruments. Each must be read for what it is — and, under this methodology, for what it cannot say.
4.1 LMArena: blind human preference — inadmissible as trust evidence
LMArena measures crowdsourced blind, head-to-head human preference: users chat with two anonymous models, vote, and votes become Elo ratings for open-ended chat preference [16]. The headline panel is Text, but separate panels exist — Agent (reported as win share), WebDev, Vision, Document, Search, Text-to-Image — and scores must not be mixed across panels [16]. Under the profile of Section 2, preference is the closest public proxy to experienced friction in open chat — people vote for what felt good to use — but it is inadmissible as evidence of low intrinsic frustration: it is gameable by fine-tuning on judge-preference data [16], and it is confounded with the sycophancy of Section 2.3, the very behavior the profile penalizes most [1]. A high preference score is therefore not reassurance; at most, a low preference score is evidence of experienced friction, ambiguous between bluntness and incompetence (Section 2.3).
Text Arena snapshot, 2026-09-22 (read live) [16]: (4) muse-spark-1.2 (xHigh) 1500 ±11; (8) muse-spark-1.3-max 1493 ±9; (9) gemini-3.8-flash-high 1493 ±9; (50) deepseek-v4-pro-high-20260813; (51) claude-sonnet-5-high; (57) deepseek-v4-pro (Elo not captured for these rows).
4.2 SimpleQA and the factuality boards: the most profile-aligned instrument, with one of the four models on it
SimpleQA's correct / incorrect / not-attempted grading is the instrument in this survey closest to Axis 1: it rewards knowing what you don't know [8]. The public boards have moved: Gemini 3.8 Flash now appears on both — #3 on the original SimpleQA board (69.4%; board updated 2026-09-06) [18] and #3 on the SimpleQA Verified board (74.6% ±2.7; board updated 2026-09-16, 47 of 50 models shown) [30]. The other three surveyed models appear on neither board. The most profile-aligned instrument therefore yields data for exactly one of the four models — a measurement, not a comparison: the other three are unmeasured here, not worse.
4.3 SWE-bench Verified and third-party coding runs: capability context with auditable harnesses
SWE-bench Verified is a 500-instance, human-filtered subset of real GitHub issues; each leaderboard entry reports % Resolved [19]. The benchmark was introduced as an academic project on resolving real-world software issues with language models [11]. Three cautions apply, each sharper under this profile. First, the top of the leaderboard is saturating: a September 2026 audit of 254 submissions across the benchmark's public splits concluded the leaderboard can no longer reliably order its top entries [12]. Second, reported scores depend on the agent scaffold and submission configuration, so figures from different runs are not directly comparable; the official leaderboard distinguishes splits and settings [19]. Third — the profile's own addition — the benchmark measures task resolution, not auditability: a resolved issue says nothing about whether the produced code was correct for the right reasons, free of hallucinated APIs, or honestly described.
Under this survey's source policy, third-party runs with published methodology are admissible as capability context. Vals AI's SWE-bench Verified run — 500 human-validated tasks, isolated Docker/cloud sandbox, minimal bash-only mini-swe-agent harness, the same system prompt for every model, provider defaults — is such a run [13]. On it, DeepSeek V4 Pro 0813 scores 96.40%, second overall, 0.60 points behind Claude Opus 5 (97.00%) [13]. No vetted third-party SWE-bench run was found for the other three surveyed models; a tracker row showing Claude Sonnet 5 at 85.2% reproduces the vendor's own figure and is therefore a vendor-sourced row, not an independent run — excluded under the policy. This survey uses the Vals figure as capability context only — never as trust evidence — and the DeepSeek case of Section 5 shows why the distinction matters.
4.4 Vals AI IOI: algorithmic code generation, all four models, one honest failure
Vals AI's IOI benchmark is a second admissible third-party run: 18 IOI problems (six each from 2024, 2025, and 2026), solved with OpenCode in an isolated C++20 sandbox, no internet, graded on hidden official tests, with the overall score as the mean of the three yearly means [14]. Scores: Gemini 3.8 Flash 56.94% | Muse Spark 1.3 Max 56.56% | DeepSeek V4 Pro 0813 51.61% | Claude Sonnet 5 45.00% (Muse Spark 1.3 non-Max: 43.94%) [14]. The run also discloses an auditable reliability signal the profile cares about: Claude Sonnet 5 exhausted its 128K output budget on 13 of 18 problems before writing code, requiring a common continuation procedure [14] — a model that cannot budget its own output is a model whose cost and behavior cannot be audited in advance. Like the SWE-bench figure, these scores are capability context, barred from the verdict by the risk-multiplier rule (Section 2).
4.5 Artificial Analysis Intelligence Index: a composite — capability context, barred from the verdict
The correct name is the Artificial Analysis Intelligence Index (often shortened to "Intelligence Index") [17]: a composite comparing models on intelligence, price, output speed, first-chunk latency, total response time, and context [17]. Live read, 2026-09-22: Gemini 3.8 Flash high 59 (pre-v4.2 scale; current-scale high not verified — the model page still shows the old-scale figure), Muse Spark 1.3 max 48 and xhigh 45, Claude Sonnet 5 38 (Adaptive Reasoning, Max Effort, current scale), DeepSeek V4 Pro 0813 max 36 (current scale) [17]. Exact global ranks were not counted — only scores — so ranks should not be inferred. Under this methodology the composite is capability context only: by the risk-multiplier rule (Section 2), a higher score without honesty data raises the stakes; it does not lower the frustration.
4.6 AA-Omniscience: calibrated abstention, all four models measured
Artificial Analysis's AA-Omniscience benchmark tests knowledge reliability under an abstention-aware design: 6,000 questions across 42 topics and six domains [29]. Its headline metric, the AA-Omniscience Index (−100 to 100), rewards correct answers, penalizes incorrect ones, and does not penalize refusal — zero means as many correct as incorrect. The companion hallucination rate is incorrect ÷ (incorrect + partial + not attempted): how often the model answers incorrectly when it should have refused or admitted not knowing. Both metrics are among the most profile-aligned in this survey: the Index prices overconfidence directly, and the hallucination rate is close to the Axis-1 unit of harm.
All four surveyed models have exact-build rows (read live 2026-09-22) [29]:
Table 2. AA-Omniscience exact-build rows (read live 2026-09-22) [29]. Index: −100 to 100; rewards correct, penalizes incorrect, refusal unpenalized. Hallucination rate: incorrect ÷ (incorrect + partial + not attempted).
| Model | AA-Omniscience Index | Hallucination rate | Reasoning effort |
|---|---|---|---|
| Gemini 3.8 Flash | 30 | 55% | high |
| Muse Spark 1.3 | 25 | 33% | max |
| Claude Sonnet 5 | 16 | 39% | max |
| DeepSeek V4 Pro 0813 | 1 | 95% | max |
Three caveats. First, reasoning effort differs: Gemini 3.8 Flash is measured at high, the other three at max — treat gaps as approximate, not a perfectly controlled comparison (this is Gap G5, Section 9). Second, the page shows no last-updated date, so no vintage is assigned. Third, Artificial Analysis operates both the benchmark and the leaderboard, and per the benchmark's paper (arXiv:2511.13029 §5), GPT-5 was used for question generation, filtering, and revision — the authors disclose possible GPT-5-family bias [29].
4.7 The FACTS suite: five factuality boards, one surveyed model
Google's FACTS benchmarks are official Kaggle leaderboards scoring factual accuracy across settings [31][32][33][34][35]. Only Gemini 3.8 Flash among the surveyed models appears on them (all read live 2026-09-22):
Table 3. Gemini 3.8 Flash on the FACTS boards (read live 2026-09-22) [31][32][33][34][35].
| Board | Score | Rank | Board updated |
|---|---|---|---|
| FACTS Suite V1 (average) | 69.4% | #2 | 2026-09-03 |
| FACTS Grounding V2 | 72.5% | #10 | 2026-09-10 |
| FACTS Search V2 | 76.3% ±1.4% | #11 of 16 | 2026-09-15 |
| FACTS Parametric V1 | 0.77 ±0.02 | #5 of 39 | 2026-09-11 |
| FACTS Multimodal V1 | 0.51 ±0.03 | #1 of 34 | 2026-09-11 |
The suite is a measurement of one model, not a comparison: the other three surveyed models have no rows and are unmeasured here, not worse. Note the Multimodal board's #1 placement coexists with a 0.51 absolute score — rank and rate answer different questions, and this profile cares about the rate.
4.8 False premises: BullshitBench V2 and InfoOps Bench
BullshitBench V2 presents 100 nonsensical or broken-premise prompts and classifies responses as clear pushback, partial challenge, accepted premise, or refusal; the metric is the share of clear pushback (refusals excluded), higher better [36]. It is the most direct Axis-2 instrument in this survey — and adjacent to sycophancy, since accepting a false premise is a form of over-agreement. Exact-build rows exist for two surveyed models (dashboard read live 2026-09-22): Claude Sonnet 5 pushes back clearly on 78.8% of prompts at max effort (80.8% at low), while DeepSeek V4 Pro 0813 does so on 35% (32% at low) [36]. Gemini 3.8 Flash and Muse Spark 1.3 have no rows; a Muse Spark 1.2 row exists but is a different build. Provenance note added on review: the BullshitBench V2 dashboard is hosted on an individual's GitHub Pages site rather than an institutional benchmark operator (Vals AI, Artificial Analysis, and the Kaggle-hosted FACTS boards, by contrast, are run by organizations with a track record and public accountability). It is included here because its prompt set and grading rubric are published and inspectable, meeting the letter of this survey's source policy — but a reader relying heavily on Axis 2 should weigh this instrument's provenance accordingly, more like a well-documented independent replication than an institutional benchmark. A second caveat: the dashboard labels the variant "Max" without spelling out "Adaptive Reasoning, Max Effort."
InfoOps Bench (arXiv:2607.28503, roster dated 2026-07-26) takes a different angle: state-backed information-operation claims across four prompt framings, grading whether the model amplifies, preserves, attenuates, or explicitly fact-checks them (judge–human agreement 92.3%) [37]. Claude Sonnet 5 posts 10.0% compliance — 90.0% integrity, rank #2 of 17 — explicitly fact-checking 79.0% of claims [37]. Caveat: the paper does not specify reasoning effort, so this is not confirmed Max Effort; and it measures propensity to support information operations, not factuality generally. The paper's DeepSeek V4 Pro row predates the 0813 build and is excluded; Gemini 3.8 Flash and Muse Spark 1.3 do not appear.
Table 4. False-premise handling (read live 2026-09-22) [36][37]. "No row" means unmeasured — not zero, not ranked.
| Model | BullshitBench V2 clear pushback (max) | InfoOps Bench integrity |
|---|---|---|
| Claude Sonnet 5 | 78.8% (low: 80.8%) | 90.0%, #2 of 17 (effort unspecified) |
| DeepSeek V4 Pro 0813 | 35% (low: 32%) | row predates 0813 — excluded |
| Gemini 3.8 Flash | no row | no row |
| Muse Spark 1.3 | no row | no row |
4.9 Humanity's Last Exam: hard factuality for all four
Humanity's Last Exam (HLE) is 2,500 expert-vetted questions across mathematics, the sciences, and the humanities, independently evaluated by Artificial Analysis as accuracy pass@1 with no tools on the text-only revision [38]. All four surveyed builds carry HLE scores:
Table 5. Humanity's Last Exam scores [25][38]. Cross-checked across AA's model article and AA-data mirrors; the official chart page was verified live.
| Model | HLE (AA, text-only, no tools) | Note |
|---|---|---|
| Gemini 3.8 Flash | 47.8% | AA-data mirrors |
| Muse Spark 1.3 | ≈47% (xhigh) / ≈49% (max) | AA article [25]; mirrors 48.7 / 49.1 |
| Claude Sonnet 5 | 41.3% | max effort; AA-data mirrors |
| DeepSeek V4 Pro 0813 | 41.0% | AA-data mirrors |
Provenance note: AA's leaderboard page renders scores in a chart; the per-model figures above are cross-checked across AA's own article and consistent AA-data mirrors (Wikipedia's HLE table citing an AA snapshot dated 2026-09-16; an AA-mirror leaderboard updated 2026-09-21). Vendor self-reports (Anthropic's 43.2, DeepSeek's own figures) are excluded as independent evidence under this survey's source policy. Muse Spark's two mirror values (48.7 vs 49.1) disagree slightly at the variant level and are reported as the range from AA's article: xhigh ≈47%, max ≈49% [25]. HLE is a hard-question accuracy test — factuality context, not an abstention-aware instrument — and enters the verdict accordingly.
5. Four Models Through the Profile Lens
Table 1a. Model standings through the profile lens (snapshot 2026-09-22).
| Ecosystem | Flagship (released) | LMArena Text (2026-09-22) | AA Intelligence Index (2026-09-22) |
|---|---|---|---|
| Anthropic (Claude) | Claude Sonnet 5 (2026-06-30) [24] | #51 claude-sonnet-5-high [16] | 38 (Adaptive Reasoning, Max Effort, current scale) [17] |
| Google (Gemini) | Gemini 3.8 Flash (2026-09-02) [23] | #9 gemini-3.8-flash-high [16] | 59 (high, pre-v4.2 scale); current-scale high not verified [17] |
| Meta (Muse Spark) | Muse Spark 1.3 (Sept 2026; exact date unverified) [25] | #4 spark-1.2 xHigh; #8 1.3-max [16] | 48 (1.3 max); 45 (1.3 xhigh) [17] |
| DeepSeek | V4 Pro 0813 (2026-08-13), MIT open weights [17] | #50 v4-pro-high-20260813; #57 v4-pro [16] | 36 (0813, Reasoning, Max Effort, current scale) [17] |
Table 1b. Price and latency (contextual — excluded from the verdict by design).
| Ecosystem | API price / 1M in-out | Output speed (AA) | First-chunk latency (AA) |
|---|---|---|---|
| Anthropic (Claude) | $2.00 / $10.00 (official) [20] | Not captured | Not captured |
| Google (Gemini) | $0.75 / $3.75 intro thru 2026-12-31; $1.50 / $7.50 from 2027-01-01 (official) [21] | 329 tok/s (3.8 Flash high) [17] | 14.26 s (3.8 Flash high) [17] |
| Meta (Muse Spark) | $1.25 / $4.25 (AA) [25] | 223 tok/s (1.3 max) [17] | 20.81 s (1.3 max) [17] |
| DeepSeek | $1.32 / $3.96 peak; half off-peak (official) [22] | Not captured | Not captured |
Google — Gemini 3.8 Flash: the best-measured factuality record, and a calibration split. Announced 2026-09-02 [23].
- The only surveyed model with exact-build scores on the abstention-aware factuality boards: #3 on both SimpleQA boards [18][30] and five FACTS rows — Suite #2, Grounding #10, Search #11 of 16, Parametric #5 of 39, Multimodal #1 of 34 [31][32][33][34][35].
- But the calibration picture splits: on AA-Omniscience it leads the four on the Index (30) while posting a 55% hallucination rate — behind Muse Spark 1.3 (33%) and Claude Sonnet 5 (39%) on the metric this profile cares about most [29], with the caveat that its effort setting (high) differs from the others (max).
- HLE 47.8% [38]; IOI 56.94%, topping the four [14].
- On false premises (BullshitBench V2, InfoOps) and sycophancy: no exact-build rows — unmeasured, not ranked.
- Under the risk-multiplier rule, its capability lead raises the stakes of its unmeasured axes rather than filling them: the most fluent model with thin honesty data is the one whose confident falsehoods would travel farthest. (Its introductory $0.75/$3.75 pricing through 2026-12-31, stepping to $1.50/$7.50 [21], is extrinsic and excluded from the verdict.)
Meta — Muse Spark 1.3: the best abstention profile, thin elsewhere. Meta's closed-frontier model, covered by Artificial Analysis in September 2026 (exact release date not verified from a vetted source) [25].
- Best preference placement of the four (#8 for 1.3-max; #4 for the 1.2 xHigh build) [16], composite scores of 48/45 [17], and 56.56% on the Vals IOI run (Max build) [14].
- On AA-Omniscience it posts the lowest hallucination rate of the four (33%) with an Index of 25 [29] — the best abstention profile in the survey; AA's own article notes the model's accuracy dip came from a higher abstention rate, which also lowered its hallucination rate [25].
- HLE ≈47% (xhigh) / ≈49% (max) [25][38].
- On the factuality boards (SimpleQA, FACTS), false premises (BullshitBench V2, InfoOps), and sycophancy: no exact-build rows — unmeasured, not ranked.
- Its intrinsic frustration is therefore not low; on three honesty axes it is unknown. (Pricing $1.25/$4.25 per 1M tokens [25] is extrinsic, excluded. Meta's vendor-authored safety report covering earlier Spark versions [arXiv:2606.12429] is excluded as independent evidence under this survey's source policy — see Section 8.)
DeepSeek — V4 Pro 0813: elite at tasks, poor at calibration — the profile's sharpest case. Released 2026-08-13 with MIT open weights [17].
- On Vals AI's documented SWE-bench Verified run it scores 96.40%, second overall [13]; on the Vals IOI run, 51.61% [14]; its composite is 36 [17].
- It sits at #50/#57 on the preference board [16] — the strongest dated signal of experienced interactive friction among the four. This is the survey's cleanest demonstration that competence and low friction are different things: a model can be second overall at resolving real GitHub issues through an audited harness and still be the most frustrating of the four to talk to.
- On the honesty axes the measurements that exist are poor: AA-Omniscience Index 1 (lowest of the four) with a 95% hallucination rate — it answers incorrectly on nearly every question it should have refused [29]; BullshitBench V2 35% clear pushback at max — it accepts the false premise more often than it challenges it [36].
- HLE 41.0% [38]. The HHEM leaderboard's 8.6% hallucination rate for "DeepSeek-V4-Pro" [15] does not qualify the 0813 build and is not treated as a measurement of it. Sycophancy: unmeasured.
- What is established is capability; what the profile demands — calibration, premise handling — measures badly where measured and is absent where not. (Official pricing is $1.32/$3.96 per 1M tokens at peak, half off-peak [22] — extrinsic, excluded.)
Anthropic — Claude Sonnet 5: the best false-premise handling among the measured, and the one auditable failure. Announced 2026-06-30 [24]; the oldest of the four.
- Sits at #51 on the preference board, the weakest single placement of the four [16], with a composite of 38 [17] and 45.00% on the Vals IOI run [14].
- The third-party record contains a disclosed reliability failure: on the Vals IOI run it exhausted its 128K output budget on 13 of 18 problems before writing code [14] — the only auditable behavioral defect attached to any of the four surveyed builds, and squarely a profile concern: unauditable effort is unauditable cost.
- But on false premises it is the best-measured model in the survey: BullshitBench V2 78.8% clear pushback at max effort (80.8% at low) [36], and InfoOps Bench 90.0% integrity — rank #2 of 17, explicitly fact-checking 79.0% of claims [37] (effort unspecified in the paper).
- AA-Omniscience: Index 16, hallucination rate 39% [29]. HLE 41.3% at max effort; GPQA Diamond 91.1% at max effort via an AA-data mirror — graduate-level science QA, factuality-adjacent [38].
- Sycophancy: no exact-build score; the nearest academic numbers test its predecessors (Sonnet 4 [26]), and family-level context is not a measurement of this model. (Official pricing $2/$10 per 1M tokens [20] — extrinsic, excluded.)
6. Verdict: What Can Be Ordered, and What Cannot
With different models measured on different axes, no honest overall ranking of the four exists: any single ordering would have to penalize models for being absent from benchmarks or invent the missing scores. This section therefore orders models within each axis, among measured models only, and states plainly where no ordering is possible. The governing rule, applied throughout: unmeasured means unranked — a model missing from a benchmark is not scored as zero and not placed below measured models.
Axis 1 — Truthfulness. Gemini 3.8 Flash is the only surveyed model measured on the abstention-aware factuality boards: #3 on both SimpleQA boards [18][30]; FACTS Suite #2, Grounding #10, Search #11 of 16, Parametric #5 of 39, Multimodal #1 of 34 [31][32][33][34][35]. No ordering is offered among the four here — one measured model is a measurement, not a comparison. On Humanity's Last Exam, all four are measured: Muse Spark 1.3 ≈49% (max) / ≈47% (xhigh), Gemini 3.8 Flash 47.8%, Claude Sonnet 5 41.3%, DeepSeek V4 Pro 0813 41.0% [25][38] — with Spark's variant values reported as the range from AA's article.
Knowledge reliability / calibrated abstention (AA-Omniscience). All four measured — the survey's only complete axis. By Index: Gemini 3.8 Flash (30) > Muse Spark 1.3 (25) > Claude Sonnet 5 (16) > DeepSeek V4 Pro 0813 (1). By hallucination rate: Muse Spark 1.3 (33%) < Claude Sonnet 5 (39%) < Gemini 3.8 Flash (55%) < DeepSeek V4 Pro 0813 (95%) [29]. The two metrics disagree at the top, and for this profile the hallucination rate is the more important one — it is the closest to the Axis-1 unit of harm. Caveat: effort settings differ (Gemini high, others max); treat gaps as approximate.
Axis 2 — Presuppositional integrity. Claude Sonnet 5 and DeepSeek V4 Pro 0813 are measured; Gemini 3.8 Flash and Muse Spark 1.3 are not — and are therefore unranked on this axis, not placed below. Among the measured: Sonnet 5 (BullshitBench V2 78.8% clear pushback at max; InfoOps 90.0% integrity, #2 of 17) ahead of DeepSeek V4 Pro 0813 (35%) [36][37].
Axis 3 — Sycophancy. No dedicated benchmark publishes exact-build scores for any of the four. No ordering is offered. This remains the survey's largest measurement gap on the honesty side — though, per Section 9.2, the design work needed to close it may already exist [45], leaving re-application rather than invention as the remaining task.
Axis 4 — Forensic code quality. No per-model auditable rates exist. No ordering is offered. Vals AI's IOI run (Gemini 3.8 Flash 56.94%, Muse Spark 1.3 Max 56.56%, DeepSeek V4 Pro 0813 51.61%, Claude Sonnet 5 45.00%) [14] and the SWE-bench Verified run (DeepSeek 96.40%, second overall) [13] remain capability context, barred from the verdict by the risk-multiplier rule — with Sonnet 5's disclosed output-budget failure [14] standing as the only auditable behavioral defect attached to any surveyed build.
The verdict. No model earns trust: every one of the four is unmeasured on at least one axis this profile treats as non-negotiable — sycophancy for all four; false premises for Gemini 3.8 Flash and Muse Spark 1.3; abstention-aware factuality for everyone but Gemini 3.8 Flash. Where measurements exist they do not converge on a single model: the lowest hallucination rate belongs to Muse Spark 1.3, the richest factuality record to Gemini 3.8 Flash, the strongest false-premise handling to Claude Sonnet 5, the strongest audited task competence to DeepSeek V4 Pro 0813 — which also posts the worst calibration numbers in the survey. That divergence is the finding: the instruments measure different things, and none of the four dominates the honesty axes. The deeper finding is about the instruments, not the models. The behaviors this user most needs measured remain the least evenly measured things in the field: truthfulness under an abstention-aware instrument exists (SimpleQA [8], AA-Omniscience [29]) but covers the four unevenly; sycophancy exists in the literature as a field-level phenomenon [1] and, increasingly, as an exact-build instrument [45] — but not yet applied to current flagships; presuppositional integrity is measured for two of the four builds and by nothing else. Third-party operators show that auditable measurement is possible — published harnesses, fixed prompts, disclosed failures [13][14][29][36] — which makes the unevenness of honesty measurement a choice by the field, not a technical impossibility. Until honesty is measured as carefully as capability, the rational posture for the engineer in this profile is uniform across all four models: verify everything, trust nothing, and treat greater capability as raising the stakes of verification rather than lowering them.
7. How to Read Benchmarks Without Fooling Yourself
1. The instrument is part of the measurement. SWE-bench Verified measures model plus harness: the official leaderboard reports each submission's scaffold and settings [19], and a September 2026 audit of 254 submissions showed entries at the top can no longer be reliably ordered [12]. Scores from different runs describe different instruments, not just different models — which is why this survey admits only third-party runs that publish their harness (Vals AI [13][14]) and excludes tracker rows that reproduce vendor figures. For the profile of Section 2, an undisclosed scaffold is not a footnote — it is a trust defect.
2. Do not expect the leaderboards to agree — they measure different things, and none measures what this profile needs. DeepSeek V4 Pro 0813 scores 96.40% on Vals AI's SWE-bench run while posting a 95% hallucination rate on AA-Omniscience [13][29]. Claude Sonnet 5 scores 38 on the composite while sitting at #51 on preference and exhausting its output budget on most IOI problems [17][16][14]. Gemini 3.8 Flash posts 59 (old scale) on the composite while sitting at #9 on preference [17][16]. These are not contradictions — preference, composite capability, scaffolded task completion, calibration, and factuality are different axes. None of them is sycophancy or auditability.
3. Watch for saturation, contamination, and gameability. When top benchmark scores cluster, further gains fall inside the instrument's noise band and the benchmark loses discriminating power — the September 2026 SWE-bench audit makes this concrete [12]. TruthfulQA faces explicit contamination and saturation critiques [7]; LMArena Elo can be inflated by training on judge preferences — which the sycophancy literature suggests may inflate agreeableness along with Elo [16][1]. Precision is not information. And for this profile, the sharpest gameability warning is directional: the better a model gets at winning preference votes, the less its preference rank says about its honesty.
4. Check provenance and date everything. Third-party benchmark data was admitted to this survey only where the operator publishes its methodology (Vals AI [13][14]; Vectara [15]; Artificial Analysis [29]; the FACTS boards [31][32][33][34][35]; BullshitBench V2 [36]) — and a tracker row that merely reproduces a vendor's own figure was excluded as a vendor-sourced row, not an independent run. Any quoted score should carry its snapshot date, its harness, and a source that meets the bar stated in Section 1. Vendor-authored evaluations of a vendor's own model are provenance failures under that bar, however rigorous their methods section. Note that "publishes its methodology" is necessary but not sufficient for trust: Section 4.8 treats BullshitBench V2 as admissible but lower-confidence than the institutional operators, precisely because publishing a rubric and having institutional accountability are different things.
5. Unmeasured means unranked. A model absent from a benchmark must not be ranked below the models on it. Absence is not a demerit, not a zero, and not an invitation to impute — it is a gap in the evidence, and the honest presentation leaves the model out of that axis's ordering entirely. Any overall ranking across axes with different coverage would have to violate this rule, which is why this survey offers none.
8. Limitations
- Snapshot. Figures date to on or before 2026-09-22; releases, prices, and scores move weekly.
- No independent experiments — and no user study. The survey synthesizes official pages read live, primary leaderboards, vetted third-party runs, and papers. The per-axis orderings in Section 6 are expert synthesis from that evidence, not measured outcomes: no users were surveyed, no friction was timed in a lab, and the ordinal positions carry no quantified gaps. Where vetted sources disagree, both figures are reported with sources, not reconciled.
- One profile, not a universal claim. The measurement methodology of Section 2 encodes a specific user's demands — absolute truthfulness, presuppositional integrity, zero sycophancy, forensic code quality. A user who values agreeableness, speed, or price would rationally reach a different verdict from the same evidence. The survey does not claim this profile is the right one; it claims only to have applied it consistently.
- Extrinsic factors excluded by design. Price, latency, and availability are reported as context (Table 1b) but do not enter the verdict — a deliberate scope choice, not an oversight.
- The honesty instrumentation is the finding. Exact-build measurements now exist but unevenly: AA-Omniscience covers all four [29]; FACTS and SimpleQA cover Gemini 3.8 Flash [18][30][31][32][33][34][35]; BullshitBench V2 and InfoOps Bench cover Claude Sonnet 5 and (BullshitBench only) DeepSeek V4 Pro 0813 [36][37]; HLE covers all four [38]. No dedicated sycophancy benchmark publishes exact-build scores for any of the four; the nearest academic numbers test older versions [26][27][28], though a validated instrument design now exists for producing such scores [45] and simply has not been re-run on current flagships.
- Could not be verified from vetted sources: vetted third-party SWE-bench runs for Gemini 3.8 Flash, Claude Sonnet 5, and Muse Spark 1.3 (only DeepSeek V4 Pro 0813 has one [13]); per-model sycophancy scores for any of the four flagships; the exact Muse Spark 1.3 release date; Artificial Analysis global rank positions (scores only); LMArena Elo values for the lower rows; Gemini 3.8 Flash's current-scale high Intelligence Index score (the model page still shows the pre-v4.2 figure); BullshitBench V2 global ranks (the dashboard showed filtered views only); the Muse Spark 1.3 HLE variant values at mirror precision (48.7 vs 49.1 — reported as the range from AA's article [25]); the Claude Sonnet 5 InfoOps Bench reasoning effort (unspecified in the paper [37]); exact-build SimpleQA/FACTS scores for Claude Sonnet 5, Muse Spark 1.3, or DeepSeek V4 Pro 0813 — several SEO-style model-comparison aggregators surface numbers for these models on these instruments, but on inspection the figures are internally inconsistent with the models' actual release dates and are not traceable to a primary run; they were checked and excluded rather than left unsearched.
- Excluded by source policy, noted for transparency: Meta's vendor-authored "Muse Spark Safety & Preparedness Report" (arXiv:2606.12429, covering earlier Spark versions); tracker rows reproducing vendor figures (e.g., the 85.2% Claude Sonnet 5 SWE-bench row); vendor self-reported benchmark figures (Anthropic's HLE 43.2 for Sonnet 5; DeepSeek's own HLE figures for V4 Pro 0813); the Vectara HHEM generic DeepSeek-V4-Pro row as a measurement of the 0813 build (it predates it [15]); news sites, social media, and individual posts as sources throughout.
- Literature scope. Findings are summarized, not methods; consult the papers for design and limitations.
9. Closing the Gaps: A Research Program and Why It Needs Funding
The unmeasured cells in Sections 5 and 6 are not noise in this survey — they are gaps in the field's instrumentation, and they are specified precisely enough to build against. This section treats them as a specification: the instruments that would close each gap, the design principles they must honor, what they would cost, and why independent funding is the mechanism that will build them.
9.1 The gaps as a specification
Table 6. The measurement gaps and what would close them.
| Gap | Scope | Why the profile needs it | Instrument that would close it |
|---|---|---|---|
| G1 Sycophancy at exact-build level | All four flagships | Axis 3; the survey's largest honesty gap | Re-run an existing validated design [45] on the four current builds — design risk is lower than previously assumed |
| G2 False-premise handling | Gemini 3.8 Flash, Muse Spark 1.3 | Axis 2 | Replicate the public BullshitBench V2 protocol on both builds |
| G3 Abstention-aware factuality | All but Gemini 3.8 Flash | Axis 1 | SimpleQA-style correct / incorrect / not-attempted grading on the three unmeasured builds |
| G4 Forensic code-quality rates | All four flagships | Axis 4 | New instrument: per-model hallucinated-API and silently-wrong-completion rates |
| G5 Effort-matched calibration | All four flagships | Cross-axis comparability | Re-run abstention-aware items at matched reasoning effort |
9.2 Four instruments, in order of cost
Instrument 1 — False-premise replication (closes G2). BullshitBench V2's protocol is public: 100 prompts, a four-way response classification, percent clear pushback as the metric [36]. Re-running it on Gemini 3.8 Flash and Muse Spark 1.3 requires no new dataset design — adapt the prompts, run the two builds, grade against the published rubric. Estimated cost: a few hundred dollars in API inference; one to two weeks of engineering. This is the cheapest gap to close and should be closed first.
Instrument 2 — Abstention-aware factuality for the three unmeasured builds (closes G3). The instrument exists; the coverage does not. Running the public SimpleQA question set (4,326 items [8]) — or its Verified successor — against Claude Sonnet 5, Muse Spark 1.3, and DeepSeek V4 Pro 0813 with the published correct / incorrect / not-attempted grading would convert a single-model measurement into a four-model comparison. Cost: a few hundred dollars in inference; grading follows the published key.
Instrument 3 — Dedicated sycophancy at exact-build level (closes G1) — now a re-application problem, not a design problem. Earlier framings of this gap treated it as needing an instrument built from scratch. That is no longer accurate. A 2026 persona-driven challenge study has already built and validated a workable design: several thousand persona-injected trials (vulnerable- and authority-framed) classified into sycophantic drift versus reversal, run against five current models with statistical testing of response patterns [45]. It found sycophantic-response rates ranging from 2.4% (Grok 4.1) to 15.3% (Gemini 3 Flash), with Claude builds clustering lowest among those tested (3.7%–5.3%). None of the five builds it covers are the four surveyed here, so it closes none of this survey's cells directly — but it removes the design-risk component from the cost estimate below. What remains is re-application: run the same protocol against Gemini 3.8 Flash, Claude Sonnet 5, Muse Spark 1.3, and DeepSeek V4 Pro 0813. Revised estimate: $2,000–$8,000 over one to two months part-time (down from the $5,000–$15,000, two-to-four-month estimate that would apply to building a new instrument), with annotation for edge-case classification — not instrument design — as the remaining cost driver.
Instrument 4 — Forensic code-quality rates (closes G4). No instrument in this survey measures what Axis 4 demands: per-model rates of hallucinated APIs, silently wrong completions, and unauditable claims about code behavior. The coding runs that exist [13][14] measure task completion, not auditability. A new instrument would pair code-generation tasks with post-hoc audits: does the emitted code import APIs that exist? Does the model's description of what the code does match what the code does? Are errors disclosed or papered over? The design is the hard part; this survey's contribution is to have specified the measurand precisely enough to build against.
9.3 Design principles the program must honor
The program must not repeat the failure modes this survey documents. Five principles, each traceable to a section above:
- Exact-build pinning. Scores attach to builds, not families (Sections 4.8, 5). A result on Sonnet 4 is not a result on Sonnet 5 [26]; a result on Gemini 3 Flash is not a result on Gemini 3.8 Flash [45].
- Preregistration. Hypotheses, items, and grading rubrics are published before models are run — the antidote to the gameability documented in Section 7.
- Open license. Prompts, rubrics, raw responses, and grades are published; harnesses are disclosed. Sections 4.3 and 4.4 show this is feasible [13][14].
- Human grading where it matters. Model-judged honesty is circular for a benchmark about honesty; the serious tier budgets human raters (Section 7).
- Unmeasured stays unranked. New instruments report per-axis scores — no composites, no imputation (Section 6).
9.4 Budget and timeline
- Replication tier: $200–$500, one to two weeks. Closes G2, and covers the inference cost of G3. One engineer, public protocols, published rubrics. Output: a dataset and a preprint appendix this survey could cite.
- Re-application tier: $2,000–$8,000, one to two months part-time. Closes G1 by re-running the validated persona-injection design [45] on the four current builds. Revised down from the original "serious tier" estimate now that instrument design is no longer the bottleneck.
- Rigorous tier: $30,000–$100,000+, team effort. Closes G4 and G5: new code-auditability instruments, effort-matched calibration, expert annotators, a multi-axis release.
Inference is not the constraint: a few thousand prompts across four models costs a few hundred dollars, and fellowship programs in this space routinely include API credits and compute support for external safety research [44]. Annotation is the constraint — which is exactly why the program needs funding rather than spare weekends.
9.5 Why this needs funding — and why it must be independent
The gaps in Table 6 are not accidents. The evidence in Section 7 points to three structural causes: the market rewards capability measurement and does not demand honesty measurement; vendors submit to the benchmarks that flatter and stay silent on the rest; and dishonesty findings are reputational risk, so labs fund the former and not the latter. No lab will pay to discover that its flagship accepts false premises most of the time. The measurement will therefore not happen inside the labs — or if it does, it will not be published.
Independence is not a preference here; it is a methodological requirement of the profile itself. Under this survey's source policy, vendor self-measurement of a vendor's own model is inadmissible as trust evidence (Section 7.4) — a honesty benchmark run by the vendor whose model it grades fails the survey's own bar. Only independent, preregistered, open-license measurement counts.
The funding exists. Safety evaluation is an explicitly funded category: the Foresight Institute's AI for Science & Safety RFP offers $30,000–$100,000 to individuals and teams producing open-source safety research, with the 2026 call closing October 31, 2026 [39]; the Long-Term Future Fund makes rapid $5,000–$100,000+ grants for technical safety work [40]; the Survival and Flourishing Fund runs quarterly rounds in the $50,000–$200,000 range [41]; Open Philanthropy's technical AI safety RFP draws on a $40M pool, beginning with a 300-word expression of interest [42]; Manifund's micro-grant regranters cover the replication tier outright [43]; and the OpenAI Safety Fellowship funds external researchers — engineers and practitioners included — to produce papers, benchmarks, and datasets in safety evaluation [44]. Given that G1 now costs less than previously estimated, it is realistically fundable at the Manifund micro-grant or Long-Term Future Fund small-grant tier rather than requiring the larger institutional grants.
The honest accounting: large institutional grants favor track records, and an independent researcher should expect to earn credibility through the tiers — replication first, then re-application. This survey is the first artifact in that sequence: the gap analysis a proposal would otherwise have to hand-wave. Sections 9.2–9.4 are the proposal in outline; written out fully, with milestones and a budget, they are submittable.
10. Conclusion
As of September 2026, the frontier is plural — and for the engineer who demands absolute truthfulness, presuppositional integrity, zero sycophancy, and forensic code quality, the evidence now supports per-axis orderings but still no overall answer. Muse Spark 1.3 has the lowest hallucination rate of the four on AA-Omniscience (33%) [29] but no exact-build rows on the factuality boards or the false-premise benchmarks. Gemini 3.8 Flash has the richest factuality record — #3 on both SimpleQA boards, five FACTS rows [18][30][31][32][33][34][35] — but the highest hallucination rate among the top three on AA-Omniscience (55%) [29] and no false-premise measurements. Claude Sonnet 5 handles false premises best among the measured (78.8% pushback; 90.0% InfoOps integrity) [36][37] but carries the weakest preference placement and the survey's only auditable behavioral defect [14][16]. DeepSeek V4 Pro 0813 pairs elite audited task competence (96.40% SWE-bench, second overall) [13] with the worst calibration in the survey (95% hallucination rate; 35% pushback) [29][36] — the proof that competence and honesty are different things. Sycophancy remains unmeasured for all four, though the instrument needed to measure it now exists and has simply not been pointed at these builds [45]. Price, latency, and availability were set aside by design: they are business facts, and a cheap, fast, confidently wrong model is not less frustrating but cheaply, quickly misleading. The verdict that matters most is the meta one: the behaviors this user most needs — not asserting falsehoods, not building on unexamined premises, not flattering, not emitting unauditable code — remain the least evenly measured by the numbers that dominate public discussion [1][2][4][6][8][29]. Third-party operators have shown that auditable measurement is possible [13][14][15][29][36]; the field has simply not yet pointed that machinery at honesty evenly. Section 9 argues that this is a funding problem with a known solution: the gaps are specified, the instruments are designed (or, for G1, already built), the budgets are modest, and the grant programs exist. The survey's final verdict is therefore not only that no model earns trust, but that trust is buildable — it just has to be measured, and measurement has to be funded. Until it is, the least-frustrating model will remain a per-axis judgment call — this survey has tried to make the judgment an informed one, and to show exactly where the information runs out.
References
- Sharma, M. et al. "Towards Understanding Sycophancy in Language Models." ICLR 2024. https://arxiv.org/abs/2310.13548
- Xiong, M. et al. "Can LLMs Express Their Uncertainty? An Empirical Evaluation of Uncertainty Expression in LLMs." ICLR 2024. https://openreview.net/forum?id=gjeQKFxFpZ
- Guo, C. et al. "On Calibration of Modern Neural Networks." ICML 2017. https://arxiv.org/abs/1706.04599
- "Beyond 'I Don't Know': A Benchmark for Uncertainty Attribution in LLMs" (UA-Bench), 2026. https://arxiv.org/abs/2604.17293v1
- Cocchieri, F. et al. "LLMs (Almost) Never Abstain Under Medical Uncertainty" (MedQAbstain). ACL 2026. https://aclanthology.org/2026.acl-long.1365/
- Lin, S. et al. "TruthfulQA: Measuring How Models Mimic Human Falsehoods." ACL 2022. https://aclanthology.org/2022.acl-long.229/
- "TruthfulQA contamination/saturation critique." arXiv 2504.17550. https://arxiv.org/pdf/2504.17550
- Wei, J. et al. "SimpleQA: Measuring Short-Form Factuality in Large Language Models." arXiv 2411.04368. https://arxiv.org/abs/2411.04368
- Li, J. et al. "HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models." EMNLP 2023. https://aclanthology.org/2023.emnlp-main.397/
- Rajpurkar, P. et al. "Know What You Don't Know: Unanswerable Questions for SQuAD." ACL 2018 (Best Short Paper). https://arxiv.org/abs/1806.03822
- Jimenez, C. E. et al. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR 2024. https://arxiv.org/abs/2310.06770
- Liu, F. et al. "Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead." arXiv 2609.17394, September 2026. https://arxiv.org/abs/2609.17394
- Vals AI — SWE-bench Verified benchmark (500 human-validated tasks, isolated sandbox, bash-only mini-swe-agent harness, same system prompt for all models), read 2026-09-22. https://vals.ai/benchmarks/swebench
- Vals AI — IOI benchmark (18 IOI problems 2024–2026, OpenCode in isolated C++20 sandbox, hidden official tests), read 2026-09-22. https://www.vals.ai/benchmarks/ioi
- Vectara — hallucination leaderboard (HHEM-2.3, >7,700 non-public documents, temperature 0), updated 2026-05-11. https://github.com/vectara/hallucination-leaderboard/
- LMArena leaderboard — Text Arena standings, read live 2026-09-22. https://lmarena.ai/leaderboard
- Artificial Analysis model leaderboard — Intelligence Index scores and latency, read live 2026-09-22. https://artificialanalysis.ai/leaderboards/models
- Kaggle — OpenAI SimpleQA leaderboard: Gemini 3.8 Flash #3 (69.4%), board updated 2026-09-06. https://www.kaggle.com/benchmarks/openai/simpleqa
- SWE-bench official leaderboards — Verified: 500 human-filtered instances; entries report % Resolved. https://www.swebench.com/
- Anthropic API pricing — Claude Sonnet 5 $2/$10 per 1M tokens, read live 2026-09-22. https://platform.claude.com/docs/en/about-claude/pricing
- Google AI — Gemini Developer API pricing: 3.8 Flash $0.75/$3.75 introductory through 2026-12-31, $1.50/$7.50 standard from 2027-01-01; read live 2026-09-22. https://ai.google.dev/gemini-api/docs/pricing
- DeepSeek API docs — Models & Pricing: V4 Pro $1.32/$3.96 peak per 1M tokens (off-peak half), read live 2026-09-22. https://api-docs.deepseek.com/quick_start/pricing
- Google blog — "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber," 2026-09-02. https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
- Anthropic News — "Introducing Claude Sonnet 5," 2026-06-30. https://www.anthropic.com/news/claude-sonnet-5
- Artificial Analysis — "Muse Spark 1.3: Meta reaches the frontier" (model article: $1.25/$4.25 per 1M token pricing, HLE xhigh ≈47% / max ≈49%, AA-Omniscience accuracy note), read 2026-09-22. https://artificialanalysis.ai/articles/muse-spark-1-3
- "The Silicon Mirror," arXiv:2604.00478 — sycophancy measured on Claude Sonnet 4 (9.6% on the TruthfulQA adversarial split) and Gemini 2.5 Flash; older versions, not the surveyed models. https://arxiv.org/abs/2604.00478
- SycEval, arXiv:2502.08177 — sycophancy evaluation covering Claude-Sonnet 3.x; older version, not the surveyed model. https://arxiv.org/abs/2502.08177
- "The Granularity Gap," arXiv:2606.05183 — tests Gemini 3.0 Flash; older version, not the surveyed model. https://arxiv.org/abs/2606.05183
- Artificial Analysis — AA-Omniscience leaderboard (6,000 questions, 42 topics; Index −100..100; hallucination rate definition), read live 2026-09-22; benchmark paper arXiv:2511.13029. https://artificialanalysis.ai/evaluations/omniscience
- Kaggle — DeepMind SimpleQA Verified leaderboard: Gemini 3.8 Flash #3 (74.6% ±2.7), board updated 2026-09-16. https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified
- Kaggle — Google FACTS Benchmark Suite V1: Gemini 3.8 Flash 69.4% (#2), board updated 2026-09-03. https://www.kaggle.com/benchmarks/google/facts
- Kaggle — Google FACTS Grounding V2: Gemini 3.8 Flash 72.5% (#10), board updated 2026-09-10. https://www.kaggle.com/benchmarks/google/facts-grounding
- Kaggle — Google FACTS Search V2: Gemini 3.8 Flash #11 of 16 (Score 76.3% ±1.4%), board updated 2026-09-15. https://www.kaggle.com/benchmarks/google/facts-search
- Kaggle — Google FACTS Parametric V1: Gemini 3.8 Flash #5 of 39 (Score 0.77 ±0.02), board updated 2026-09-11. https://www.kaggle.com/benchmarks/google/facts-parametric
- Kaggle — Google FACTS Multimodal V1: Gemini 3.8 Flash #1 of 34 (Score 0.51 ±0.03), board updated 2026-09-11. https://www.kaggle.com/benchmarks/google/facts-multimodal
- BullshitBench V2 — false-premise benchmark (100 prompts; % clear pushback): Claude Sonnet 5 Max 78.8% / Low 80.8%; DeepSeek V4 Pro 0813 Max 35% / Low 32%; dashboard read live 2026-09-22; hosted on an individual's GitHub Pages site rather than an institutional operator (see Section 4.8). https://petergpt.github.io/bullshit-benchmark/viewer/index.next.html
- InfoOps Bench, arXiv:2607.28503 (roster 2026-07-26): Claude Sonnet 5 compliance 10.0% / integrity 90.0% (#2 of 17), explicitly fact-checked 79.0%; reasoning effort unspecified. https://arxiv.org/pdf/2607.28503
- Artificial Analysis — Humanity's Last Exam leaderboard (text-only revision, pass@1, no tools), chart page verified live 2026-09-22; per-model figures cross-checked across AA's article and AA-data mirrors. https://artificialanalysis.ai/evaluations/humanitys-last-exam
- Foresight Institute — AI for Science & Safety Nodes Request for Proposals 2026: $30,000–$100,000 for open-source AI safety research; individuals, teams, and organizations eligible; application deadline 2026-10-31. https://foresight.org/grants/ai-science-safety-nodes-rfp/
- Long-Term Future Fund — rapid-turnaround grants ($5K–$100K+) for technical AI safety research, compute, and stipends. https://funds.effectivealtruism.org/funds/long-term-future
- Survival and Flourishing Fund — quarterly grant rounds for AI safety and catastrophic-risk research. https://survivalandflourishing.fund/
- Coefficient Giving (Open Philanthropy) — Request for Proposals: technical AI safety research ($40M pool; 300-word expression of interest). https://coefficientgiving.org/funds/navigating-transformative-ai/request-for-proposals-technical-ai-safety-research/
- Manifund — micro-grants via community regranters for exploratory AI safety projects and open tooling. https://manifund.org/
- OpenAI Safety Fellowship — paid fellowship funding external researchers (Sept 2026–Feb 2027 cohort); safety evaluation a priority area; stipend, API credits, and mentorship; expected output a paper, benchmark, or dataset. https://www.helpnetsecurity.com/2026/04/07/openai-safety-fellowship-applications/
- "Evaluating Sycophancy in Frontier Models Using Persona-Driven Challenge." medRxiv, May 2026 — persona-injection sycophancy design (vulnerable/authority framings, drift-vs-reversal classification) scoring exact builds of Gemini 3 Flash (15.3%), GPT-5.4 (8.8%), Claude Opus 4.6 (5.3%), Claude Sonnet 4.6 (3.7%), and Grok 4.1 (2.4%); does not cover the four models surveyed here, but validates the instrument design (see Section 9.2). https://www.medrxiv.org/content/10.64898/2026.05.17.26353406.full.pdf