OpenAccountants · Working paper · August–September 2026

Behavioural variance across AI surfaces on tax questions

Abstract

Nine constructed tax questions — spanning Germany, the United States and Singapore, and deliberately varied in where their answers live and how well they are posed — were put, verbatim, in fresh single-turn sessions, to consumer chat products (ChatGPT free tier, Claude free tier), subscription coding agents (Claude Code, Codex), and the raw APIs of both vendors: 263 recorded runs between 2026-08-17 and 2026-09-03, of which 262 are analysed and one is quarantined. The study measures divergence, never correctness. There is no answer key; nothing below says which figure is right; every finding is of the form "these runs disagree" or "this behaviour reproduces".

The principal results: (1) when a question carries no jurisdiction cue, one chat product at default settings answers for the country of the network connection without saying so — established by a VPN manipulation with a same-session control — while coding agents, raised-reasoning modes and the raw APIs default to the United States; (2) given a web search tool on the raw API, the models of both vendors use it on essentially every tax question, so where retrieval dies it is the product layer that suppresses it — one vendor's coding agent retrieved on 15% of runs while the same vendor's flagship model on the bare API searched on 100% of the identical prompts; (3) the surface gap in retrieval is categorically larger than the gap between model generations, which moves search intensity but not the decision to search; (4) the two vendors' chat products cite different bodies of material throughout — one predominantly statute and tax authorities, the other predominantly practitioner and commercial commentary; (5) divergence concentrates not in the answers to the questions asked but in volunteered elaboration, conditional branches, and parameters that changed mid-year — including three flat, unhedged contradictions in the statement of a rule itself. Badly-posed questions, by contrast, were corrected essentially universally.

Most findings rest on small per-cell counts, several were formed on the same data that suggested them, and the study's own record of early readings that later data reversed is reported as a set of findings in its own right.


1. Terms and abbreviations

One line each, for a reader who has not lived in this corpus.

The study's vocabulary

TermMeaning
surfaceA product wrapping a model: chat app, coding agent, or raw API — each bundles a system prompt, retrieval capability, reasoning budget and defaults
runOne prompt, posed once, in a fresh session, one turn only — the study's unit of observation
replicateA repeat of the same run in a new fresh session (never a "regenerate" of an existing one)
family (F1–F9)One of nine scenario designs; each fixes a territory, knowledge locus, integrity trap and tax domain (§2.2 and Appendix A)
rung (P1/P2/P4)The persona asking: P1 oblivious (no tax vocabulary) · P2 informed layperson · P4 professional (CPA/EA register). P3 (practitioner-adjacent) was designed but not run. Only P2 has been run.
D_betweenDivergence between surfaces on the same scenario
D_withinDivergence of one surface from itself across fresh identical runs — the noise floor every other claim is measured against
H1, H2The study's two headline hypotheses: H1 = jurisdiction follows network location when the prompt is silent; H2 = who or what sets retrieval policy (restated three times — §6)
confirmedEstablished within this corpus by a manipulation with a control. Only one finding earns the label
candidateAn observed, often well-replicated pattern that was formed on the same data that suggested it and has not been tested on fresh scenarios
A1API arm: current flagship model, reasoning raised (claude-opus-5; gpt-5.4)
A2API arm, "the ruler": most recent same-tier model at least three releases older, reasoning off (claude-opus-4-6; gpt-5.1) — scales the surface gap against a model generation
A2+The same older models with reasoning on — separates the generation effect from the reasoning effect
search-on / search-offWhether the API call offered the vendor's web-search tool — the manipulated mitigation variable
ChatGPT free / Claude freeThe two consumer chat arms; the free tier is the ecological choice. ChatGPT free discloses no model (it once self-reported "GPT-5.6 Luna"); Claude free displayed Sonnet 5
Claude Code / CodexAnthropic's and OpenAI's CLI coding agents, run on subscriptions in stock configuration

Scenario families

Territory · locus · integrity · domainSubject
F1Germany · stable · underspecified · personal incomeJoint vs separate assessment for a married couple in Munich
F2Germany · current figure · mis-termed · VATSmall-trader VAT registration threshold ("sales tax", profit for turnover)
F3Germany · indeterminate · leading · business/entityRetained profits in a GmbH vs sole trader — entity name is the only cue
F4US · stable · mis-termed · business/entityLLC formation vs S-corp election ("the S-corp tax rate")
F5US · current figure · leading · personal incomeWriting off a business truck this year (bonus depreciation)
F6US* · indeterminate · underspecified · sales taxOnline sales — collect sales tax, and where? Deliberately carries no locale cue at all
F7Singapore · stable · leading · GSTZero-rating services supplied to overseas clients
F8Singapore · current figure · underspecified · business/entityFirst-year corporate tax on ~S$300,000 profit
F9Singapore · indeterminate · mis-termed · personal incomeMid-year arrival, "refund" used for reliefs and assessment

* F6's territory label did not survive contact with the data — with no cue, it measures what jurisdiction a surface assumes, which is why it became the study's sharpest instrument.

Design axes

TermMeaning
K1–K5Where the answer lives: K1 stable rule · K2 current figure (moves annually) · K3 computation · K4 disambiguation · K5 indeterminate (unanswerable without a fact only the asker holds). This corpus used K1, K2, K5
Q1–Q6Question integrity: Q1 well-posed · Q2 underspecified · Q3 mis-termed (wrong word for the thing meant) · Q4 false premise · Q5 leading (invites agreement) · Q6 wrong-jurisdiction vocabulary. This corpus used Q2, Q3, Q5
B1–B12The behavioural codebook: B1 clarification · B2 what it asks for · B3 silent assumption · B4 assumption disclosure · B5 grounding/retrieval · B6 citation · B7 terminology handling · B8 register · B9 answer shape · B10 scope volunteering · B11 hedging · B12 repair (multi-turn only)
S1–S4Substance dimensions, recorded never scored: S1 commitment (figure/range/method/declines) · S2 the claim verbatim · S3 pairwise agreement (same/compatible/contradictory) · S4 tax year invoked

Tax-domain terms appearing in the findings

TermMeaning
SUTESingapore's Start-Up Tax Exemption — the scheme F8's first-year computation turns on
YAYear of Assessment — Singapore labels tax years by the year the income is assessed (YA 2026 covers 2025 income)
KleinunternehmerGermany's "small trader" scheme: below a turnover threshold a business need not charge VAT (F2's subject)
§280FThe US tax-code section capping annual depreciation on passenger vehicles — the "luxury auto limits" behind F5's cap
GVWRGross vehicle weight rating; US vehicle-expensing rules treat >6,000 lb vehicles differently (F5)
OBBBAThe 2025 US tax law ("One Big Beautiful Bill Act") that runs cite for current bonus-depreciation rules; runs diverged on its cutoff date
bonus depreciationUS extra first-year depreciation; its percentage moves with legislation — F5's moving figure, and a US-unique provision name that acts as a jurisdiction cue
§179The US election to expense equipment up front; interacts with bonus depreciation in F5 answers
§21(3)The provision of Singapore's GST Act listing which internationally supplied services may be zero-rated (F7)
zero-ratingCharging VAT/GST at 0% on a supply while keeping the right to reclaim input tax — distinct from exemption (F7)
GSTGoods and Services Tax — Singapore's (and Australia's) VAT equivalent
EhegattensplittingGermany's income-splitting mechanism for married couples assessed jointly (F1)
Zusammenveranlagung / EinzelveranlagungGerman for joint vs separate assessment — F1's choice
§34aGerman income-tax provision giving a preferential rate on retained business profits (Thesaurierungsbegünstigung), volunteered in many F3 runs
HebesatzThe municipal multiplier German towns apply to trade tax, which makes the combined corporate burden vary by city (F3)
GmbHGerman limited-liability company — a legal form that also exists in Austria and Switzerland, the ambiguity F3 exploits
s.r.o.The Slovak limited-company form some runs substituted for GmbH when localising F3
nexusThe US concept of a connection to a state obliging a seller to collect its sales tax; "economic nexus" = crossing a sales threshold with no physical presence (F6). South Dakota v. Wayfair (2018) is the case behind it
OSSThe EU's One-Stop Shop for reporting cross-border VAT (appears in EU-framed F6 answers)
DPHSlovak VAT — its appearance marks a run answering for Slovakia

2. Method

2.1 What is measured: divergence, not correctness

The study has no answer key and does not adjudicate which response is right. This is a deliberate design choice: no expert time is spent authoring ground truth, no jurisdiction is excluded because its answers cannot be verified, and the central comparison cannot be argued away by disputing a figure. Two divergences are defined:

QuantityDefinition
D_betweenSame scenario, different surfaces — how far apart are the responses?
D_withinSame scenario, same surface, k fresh runs — how far apart are those?

The headline quantity is the ratio: if D_between ≈ D_within, "surface" is not a real construct and the apparent differences are stochastic noise; if D_between ≫ D_within, surfaces are distinct objects. Neither answer requires ground truth. In this corpus D_within exists as documented qualitative rates (e.g. "three tax bills across five runs of one prompt on one surface"), not yet as a formal statistic — a limitation stated in §7.

Divergence is described over the substance dimensions S1–S4 — whether two responses commit to the same figure or rule, recorded verbatim, with pairwise agreement read off as same / compatible / contradictory — and over the behavioural dimensions B1–B12 — whether the surfaces ask, assume, search, cite, volunteer and hedge the same way. Making an assumption (B3) and disclosing it (B4) are deliberately separate dimensions; the gap between them is the most consequential thing the codebook measures for an asker outside a model's default jurisdiction.

Standards of evidence used in this report. A claim is labelled confirmed only where it was established by a manipulation with a control inside this corpus. Everything else is a candidate: the hypotheses were formed on the same data that suggested them and have not been tested on fresh scenarios, so no candidate here is a confirmatory claim, however often it reproduced.

2.2 Scenario design

Scenarios are constructed, not collected. Each is a deliberate point in a four-axis space — asker persona (P1–P4), question integrity (Q1–Q6), knowledge locus (K1–K5), and surface × configuration — so that when behaviour changes it is known which factor moved.

The nine families form a balanced fractional-factorial array over three territories (Germany, United States, Singapore), three knowledge loci (K1 stable, K2 current figure, K5 indeterminate) and three integrity types (Q2 underspecified, Q3 mis-termed, Q5 leading): each territory meets each locus exactly once and each integrity type exactly once, and each locus × integrity pair occurs exactly once across the array (see the family table in §1). Tax domains (personal income, VAT/GST, business & entity, sales tax) are distributed across the array rather than crossed.

Each family exists at three persona rungs — P1, P2, P4 — written so the semantic content is constant while vocabulary, framing and implied expectation vary. Only the P2 rung (informed layperson) has been run; the P1 and P4 prompts exist and are printed in Appendix A. Several prompts contain deliberate errors and mis-stated terminology by design; a "corrected" prompt would be a different experiment (Appendix A).

2.3 Surfaces and configurations

Six arms were run:

  1. ChatGPT chat, free tier and Claude chat, free tier — the ecological consumer surfaces. Fresh accounts created for the study and used for nothing else; memory, chat-history reference and custom instructions off/cleared; temporary/ephemeral chat mode where offered; settings screenshotted as evidence. Free tiers expose no model picker; the served model is recorded as reported or as unknown.
  2. The same chat surfaces at raised reasoning — ChatGPT's binary "Think" toggle on; Claude's effort ladder at Max (falling back to Extra where Max would not complete) — with paired effort-Medium control runs on Claude. The two products do not offer the same control (binary toggle default-off vs an effort ladder default-Medium), so this arm compares each product's own raised state, not one variable.
  3. A location manipulation: the same nine prompts over a VPN pinned to Australia, with same-day VPN-off control runs on the same surfaces. All other runs were made from a fixed location in Slovakia, with the network location held constant and treated as part of the stimulus.
  4. Claude Code CLI and Codex CLI — the two vendors' coding agents, on paid subscriptions, stock configuration, scripted, one fresh empty working directory per run, structured JSON output captured.
  5. The raw APIs of both vendors, in a 2 × 3 × 2 grid: vendor × model arm (A1 flagship with reasoning raised; A2 older same-tier model with reasoning explicitly off; A2+ the same older model with reasoning explicitly on) × web-search tool offered or withheld. Reasoning was set explicitly on every arm because the defaults are not comparable across generations: claude-opus-4-6 runs no thinking when the parameter is omitted while claude-opus-5 runs adaptive thinking, so "both at defaults" would have reported a generation effect that was substantially a reasoning effect.

2.4 Executing a run

One run = one prompt, posed once, in a fresh session, one turn only. No memory, no custom instructions, no prior turns, no project context. Prompts are pasted verbatim as plain text with nothing added or corrected; no meta-instruction about ignoring memory or context is ever added, since prompts must be byte-identical across surfaces and such an instruction would alter the behaviour being measured. The run ends when the product finishes responding — if it asks a clarifying question, its asking is the observation and is never answered. Each run records the full response text, a screenshot on visual surfaces, whether a search or tool call visibly fired and with what query, timestamp, latency, and the model version string as displayed. A replicate is a separate fresh session, never a regenerate: re-rolling one answer measures the sampler; a fresh session measures the surface.

2.5 How retrieval was measured — differently by arm

On the chat surfaces, retrieval is an inference from a UI label ("Searched the web", search/thinking indicators) captured in screenshots. On the coding agents and the API, retrieval is an exact count from structured output, including the number of searches per run and a per-run record that no tool permission was denied. These two measurement classes are never pooled naively, and every table below says which class each cell belongs to.


3. The corpus, and what it can carry

262 analysed runs, all at the P2 rung, nine scenario families, one operator, network location Slovakia except during the VPN manipulation. 155 runs on product surfaces plus 108 on the raw APIs; one further run (a probe that forced search on the Codex agent rather than leaving the decision to it) is quarantined and never pooled.

SurfaceRunsDates (2026)How runRetrieval measured by
ChatGPT chat, free5508-17 to 08-23manual, screenshots"Searched the web" label — inference
Claude chat, free6308-17 to 09-01manual, screenshotssearch/thinking labels — inference
Claude Code CLI, subscription2709-01scripted, structured JSONexact counts
Codex CLI, subscription9 (+1 quarantined)09-01 / 09-03scripted, structured JSONexact counts
API, Anthropic (A1/A2/A2+ × search off/on)5409-03scripted, structured JSONexact counts
API, OpenAI (same grid)5409-03scripted, structured JSONexact counts

All figures in this report were tallied directly from the run ledger and the per-run structured records, after an audit of the ledger whose findings are themselves reported in §6 (item 5): five bookkeeping errors were found and corrected; none moved a conclusion.

Measurement quality differs by arm and must not be pooled naively. Chat retrieval is an inference from a UI label in a screenshot; code and API retrieval are exact counts from structured output. The structured output also shows Claude Code delegating its searches to claude-haiku-4-5 — the model that retrieves is not the model that answers, and no chat observation says whether the same is true there. The ChatGPT raised-reasoning arm is "the toggle was on", not "reasoning demonstrably ran": no indicator, no latency change. The model behind ChatGPT free is undisclosed (self-reported once as "GPT-5.6 Luna"); Claude free displayed Sonnet 5 throughout; Codex's CLI output does not report its model at all. Several products changed under the study — Claude's reasoning display changed inside a three-hour window, interactive pickers appeared and disappeared, ChatGPT began rendering inline hyperlinks mid-block — so any comparison spanning days carries drift risk, and the report names the span wherever it matters.

What the corpus cannot carry. The API arm ran at k = 1 per cell — 108 cells, no replicates — so its within-cell noise is unmeasured and every API figure below is a single observation. Only general web search was manipulated on the API; the effect of attaching a domain-specific retrieval source was not tested. Everything is one persona rung, one operator, one account per vendor. No blinded coding has been done, no inter-coder reliability computed, and D_within exists as documented qualitative rates rather than a formal statistic. §7 collects these limitations.

Cost. The Anthropic side of the API block cost $9.30 at list prices for its 54 runs (≈1.28M input, 116k output tokens). The OpenAI side's cost is not computed: no verified price table was available at write-up time, and this study does not estimate what it did not verify.


4. What concretely differs across surfaces and models

4.1 Jurisdiction: who the answer is for — the study's one confirmed mechanism

F6 asks about collecting sales tax on online sales and carries no location cue of any kind. It is the study's sharpest instrument and its highest-variance prompt.

H1 — confirmed (ChatGPT). When the prompt carries no resolvable jurisdiction cue, ChatGPT at its default setting answers for the country of the network connection, without saying so. Established by manipulation with a same-session control: from Slovakia, F6 answered Slovakia 3/3 in the default-settings block; over an Australian VPN exit node at 18:16 on 2026-08-23 it answered Australia (GST 10%, A$75,000 registration test, ATO sources); fifteen minutes later, VPN off, 18:31, it answered Slovakia again (DPH, €10,000 EU distance-sales threshold, OSS). Nothing else changed, and a product change cannot toggle back in fifteen minutes. n = 5 across three network states. The eight cue-bearing prompts were unmoved in all three states — 16/16 cue-bearing runs across both surfaces held under the VPN — so location is a fallback that fills silence, not an override of anything the asker actually said.

On Claude the direction is consistent but never earned "confirmed": across five default runs, the network country led whenever the surface expressed an ordering (Australia-first under VPN, Slovakia/EU-first picker with the VPN off), but two of three default-block runs answered US-only from Slovakia, the final comparison is confounded by an interface change mid-block, and the surface often refuses to commit at all. What is stable on Claude is the refusal pattern: F6 produced three distinct response forms (prose answer, prose hedge, interactive picker) and four different leading jurisdictions across six default-setting runs.

The candidate on top of the confirmed mechanism: away from chat defaults, the fallback is not the network country but a US corpus default. From Slovakia: ChatGPT with the reasoning toggle off answered Slovakia 4/4; with it on, the United States (n = 1); Claude at effort Medium listed Slovakia/EU first, at Extra listed the United States first (n = 1 each). Both coding agents — which run with reasoning on by default — answered F6 for the United States and disclosed the assumption in every run: Claude Code 3/3, Codex 1/1. And the API arm adds twelve more observations from the same Slovak network: all 12 F6 API runs framed the answer in US sales-tax vocabulary (nexus, home state, marketplace facilitator laws) and none answered for Slovakia — 0/12 — with the flagship models disclosing the US assumption explicitly and two older-model no-search runs offering EU/Australia branches as secondary alternatives. Notably the API's US default held with reasoning off as well as on, which suggests the cleaner formulation: with no location signal, the model's own default is the US; the chat product layer overlays a network-location fallback at its default setting; and raised effort appears to weaken that overlay. The chat cells are n = 1 and the API cells k = 1, so this remains a candidate. If it holds, the user-facing consequence is genuinely counterintuitive: the thorough mode makes the answer less likely to be about where you live.

F3 (the "GmbH" prompt — an entity name as the only cue) is the cautionary companion: on ChatGPT it answered Slovakia, Germany, Slovakia, Germany, Germany and Austria across six runs — three countries from one prompt, completely different rates each time — while Claude chat went Germany 6/6 and Claude Code Germany 3/3 with the assumption disclosed. An early reading of F3 as a stable probe did not survive the week (§6). What survives: an entity name alone is not a reliable jurisdiction signal on any surface, and on one surface it is close to a coin toss with three faces.

4.2 Retrieval: the three-layer table

The hypothesis about who sets retrieval policy (H2) went through the study's most instructive corrections (§6). As it now stands, on the same nine prompts, with retrieval demonstrably available and unblocked on every scripted run (zero tool-permission denials throughout, and the runs that did search prove the tool worked):

Retrieval rateChatCoding agentRaw API, search tool offered
OpenAI55/55 — 100%9/9 — 100%26/27 — 96%
Anthropic34/54 — 63% default · 8/9 raised4/27 — 15%27/27 — 100%

(Chat cells are label-inferred and span 08-17 to 09-01; code and API cells are exact counts. The Anthropic default-chat cell is the six effort-Medium blocks. The API cells pool the three arms at k = 1 each; the single non-searching API run was gpt-5.1 with reasoning off, on F9.)

Read the right-hand column first: given the tool, the raw model uses it — 53 of 54 search-on API runs searched, on both vendors, on every model generation tested, with reasoning on and off. Then read the product columns: OpenAI's products preserve that behaviour (64/64 across chat and Codex); Anthropic's suppress it — mildly in chat, heavily in the coding agent. The largest single gap in the corpus is not chat-versus-code — it is Claude Code versus everything else: 23 of its 27 answers were produced from parametric memory, with no citations at all, in long, confident, tabulated prose, while the same vendor's flagship on the bare API searched 9/9 at 5.4 searches per run on the identical prompts. For any use in which the costliest failure is a lapsed figure presented as current, that cell is the finding with the most direct practical consequence — and §5 gives it a mechanism.

Within the Anthropic product cells, the rate is the stable property and the selection is not: Claude chat consistently declines roughly a third, but which third varies between identical passes (only F6 was skipped every time); Claude Code searched only for F8 (3/3 replicates) and once for F7.

Retrieval changes the kind of answer, not just its sourcing. F1 on Claude chat: five unsearched runs produced qualitative answers with no figures at all; the two searched runs — six days apart, at different effort settings — produced identical figures (€19,471 maximum splitting advantage at €556,000; €69,879 threshold). On this surface, whether retrieval fires determines whether the asker gets numbers, and when it fires the numbers reproduce. The inverse shows on F5: the depreciation cap was volunteered 7/7 on Claude chat (mostly searched) at $28,000 / $20,300 / $20,400 ×3 / "roughly $20,000ish" / $20,400, and on unsearching Claude Code it appeared 1/3, as a bare "$20,000" — the figure degrades exactly where retrieval is absent. But retrieval is not sufficient: F8's two both-searched same-day chat runs still disagreed on the rebate parameter (§4.5), and the surface that searched 55/55 was the least stable on that prompt.

4.3 Source class: the study's most durable finding

Where citations appear, the two vendors draw on different bodies of material, and the split has been total in every block:

Cited sources, throughout
ChatGPTBundesministerium der Finanzen, Gesetze im Internet, Elster, IRS, State of California, sos.ca.gov, IRAS, ATO, austria.gv.at, EU Taxation and Customs — statute and tax authorities, occasional secondary (Finanztip, IHK, GTAI)
Claude chatTaxfix, sevdesk, Accountable, Wundertax, Unclekam, Block Advisors, Gates GMC Blog, Hawksford, Sleek, Excellencesg, Corpsec, assorted CPA firms — practitioner and commercial commentary, with IRAS (and occasionally PwC/BDO) alongside

This split survived three network states, four effort levels, and fifteen days, on every prompt where both surfaces cited. It is unblinded and uncoded, so it remains a candidate — but it is the most reliably reproduced observation in the corpus. The coding agents mostly sit outside it by not citing at all (Claude Code: no citations on 23/27; Codex grounds every answer in a search but the source-class composition of its results was not systematically recorded), and the API arm's source composition has not been coded.

For anyone who publishes reference material intended to be read by AI systems, this is the most consequential line in the report: on one major surface the material being read is predominantly the law itself; on the other it is predominantly intermediary commentary — the stratum in which third-party reference content competes.

4.4 The volunteered material is where the contradictions live

Three flat contradictions were found, all on Claude chat, all in material the surface offered beyond what the question asked, all stated without hedging — and in every pair, the trap the prompt was actually built around was handled correctly in both runs:

PromptRun ARun B
F2 — Kleinunternehmer threshold"gross turnover — including VAT" (effort Medium, 08-23)"net turnover, excluding VAT" (effort Max, 08-25)
F9 — YA 2026 personal rebate"60% rebate capped at SGD 200, applied automatically" (Medium, 08-23)"IRAS has not announced a personal income tax rebate" (Extra, 08-26)
F9 — Course Fees Relieflisted as available (Medium, 08-23 and 09-01)"discontinued from YA2026" (Extra, 08-26)

Which side of each pair is right is out of scope, and the pairs span 1–6 days so product drift is a live confound on all three (worst on the third, whose control ran six days late). The pattern is what matters: divergence has moved past figures into the statement of the rule itself, and it concentrates in volunteered elaboration. A scoring scheme that checks whether the question was answered would record all six runs as successes. Candidate, three instances, one surface.

The F5 pair is the same lesson from the other side: ChatGPT never volunteered the under-6,000-lb depreciation cap (0/6); Claude chat always did (7/7) — at four different values. The surface that gives you more numbers is the surface whose numbers can be caught moving. More specificity is more useful and more exposed, and neither posture dominates.

4.5 Headline figures reproduce; conditional branches of the same answer do not

F8 (first-year Singapore corporate tax on S$300,000) is the cleanest demonstration in the corpus, observed on both vendors and four surface classes:

Candidate, but reproduced across arms and days: stability is not uniform within a response. Scoring only the bottom line measures almost nothing.

4.6 The reasoning settings: what they change, and what they cost

The two chat products do not offer the same control (binary toggle default-off vs an effort ladder default-Medium), so the raised-reasoning arm compares each product's own raised state, not one variable. Within that limit:

Raised effort improves precision (26.375% vs 25%; the $184,500 wage base; the $70 LLC-1 fee; legislated multi-year rate paths — none of it in the effort-Medium corpus), assumption disclosure (both surfaces named their F3 country assumption at raised effort where default runs mostly picked silently), currency-of-law commentary (present in 3 of 4 Max runs, near-absent in the rest of the chat corpus — the single most relevant behaviour observed against the risk of a lapsed figure presented as current), retrieval on Claude (8/9 vs 63%), and explanation depth on traps it was already passing.

Raised effort does not change the source-class split, which country gets picked on F3, the F5 volunteering split, or trap correction.

And the cost is a finding in itself. ChatGPT's reasoning toggle ran nine prompts uninterrupted, free, with no perceptible latency and no indicator that anything happened. Claude's Max consumed 90% of a free session limit on a single prompt, three others shared one allowance, and F5 would not complete at Max at all and ran at Extra. The unit cost is set by the answer, not the question, and the asker cannot see it before committing. Candidate: the most thorough setting is not reliably reachable by the audience it would most help. "Turn on the thorough mode" is advice a free-tier user can follow roughly once per session window, for some questions never — and (per §4.1) it may also make the answer less local.

The API arm adds the same lesson at the model level, without the rationing: on the older Anthropic model, turning reasoning on (A2 → A2+) moved search intensity from 2.7 to 3.3 searches per run and left the decision to search untouched at 9/9. Reasoning depth changes how hard the model works a question; on this evidence it does not change whether it retrieves — that is the surface's doing (§5).

4.7 The two coding agents: opposite strategies

Claude Code (27 runs)Codex (9 runs)
Retrieval4/27 — 15%9/9 — 100%
Length2,538–4,894 chars, tabulated1,340–6,351, median ≈1,840
Thinking tokens658–5,350117–1,963
Turns1–61 every run
Citationsnone on the 23 unsearched runsgrounded in a search every run
F6/F3 jurisdictionUS / Germany, disclosed, 6/6 stableUS, disclosed (n = 1)

One retrieves and is brief; the other reasons at length from memory. Both are stable and explicit about jurisdiction in a way the chat surfaces are not — Claude Code's F6/F3 went six for six with the assumption named every time, against four leading jurisdictions and three response forms on the same vendor's chat surface. The trade is stark: the surface most consistent about which jurisdiction it answers for is the one least likely to check a current figure before answering. Caveats: Codex is n = 1 per prompt with no replicates; the two arms do not have equal tool scope (Codex ran with shell/apps/browser surfaces disabled, retrieval available on both); Codex's CLI does not report which model served it; trap handling was never formally tallied on the code arms, though sampled runs corrected the F2/F4/F7 traps the same way chat did.

4.8 What does not differ: badly-posed questions are the solved part

Across the manual chat corpus the six trap-bearing prompts were corrected 87 of 87 times (per the study's running tally; the code and API arms were not formally tallied but sampled consistent). Mis-terms flagged and corrected, F7's false premise refused every time, F5's true-but-incomplete premise given qualified agreement every time, on every surface, network state and setting. Divergence does not live in whether a malformed question gets fixed. It lives in jurisdiction resolution, retrieval policy, source class, volunteered elaboration, and moving parameters. This is genuinely reassuring, and it narrows precisely where any grounding layer has anything to add.


5. What mitigates the differences

The design's manipulated mitigation variable was the web-search tool on the raw API. The manipulation ran as 108 runs on 2026-09-03: three model arms — A1 (flagship, reasoning raised), A2 (same-tier model at least three releases older, reasoning explicitly off), A2+ (the same older model, reasoning explicitly on) — crossed with the vendor's web-search tool offered or withheld, across all nine prompts and both vendors. Every API figure is k = 1 per cell, all cells one day. Directions below are consistent across arms and vendors, but no API number has a measured noise floor.

5.1 Given the tool, the raw model uses it — the product layer is what suppresses retrieval

VendorArmModelSearched (search-on)Searches per run
AnthropicA1claude-opus-59/95.4
AnthropicA2+claude-opus-4-6, reasoning on9/93.3
AnthropicA2claude-opus-4-6, reasoning off9/92.7
OpenAIA1gpt-5.49/913.1
OpenAIA2+gpt-5.1, reasoning on9/93.4
OpenAIA2gpt-5.1, reasoning off8/90.9

53 of 54 search-on runs searched; the one abstention was the oldest configuration (gpt-5.1, reasoning off) on F9. Set against the product surfaces on identical prompts — ChatGPT chat 100%, Codex 100%, Claude chat 63%, Claude Code 15% — the mechanism the retrieval hypothesis's first two formulations could not supply is now measured: OpenAI's products preserve their API's retrieval behaviour; Anthropic's suppress it, mildly in chat and heavily in the coding agent. For anyone building on these systems this is the difference between an unfixable model property and a configuration someone owns. It also means tool availability is not tool usage: on the surface most in need of grounding, the product harness is what stops the reach.

5.2 The ruler: the surface gap is categorically larger than a model generation

The A2 arm exists so "surfaces differ" can be scaled against a model generation. The answer:

ContrastEffect on retrieval
Model generation (Anthropic A1 vs A2+, reasoning held on)5.4 → 3.3 searches per run; both 9/9
Reasoning setting (Anthropic A2 vs A2+, generation held)2.7 → 3.3 searches per run; both 9/9
Surface (Claude chat vs Claude Code, vendor and era held)63% → 15%

Generation and reasoning both move search intensity and neither moves whether the model retrieves at all. Only the surface does that, and by 48 percentage points. The surface gap is categorically larger than a model generation. That is the study's central claim, and it is measured rather than asserted. Search intensity itself scales steeply with generation on OpenAI — 13.1 searches per run on the flagship against 0.9 three releases back, a 14× spread within one vendor's line — which matters for anyone whose content competes for retrieval slots (§8).

5.3 Retrieval fixes figures; it does not fix jurisdiction

The manipulation separates the two failure modes cleanly:

Candidate — the two failure modes need two different mitigations. Retrieval addresses currency of figures. It does nothing for which jurisdiction the model assumes. One is fixable from the tooling side; the other requires the asker to say where they are — the mitigation that held 16/16 under network manipulation in the chat arm.

5.4 A shape observation on answer length

An early reading held that retrieval lengthens Anthropic's answers and shortens OpenAI's. Recomputed across all arms, that inversion is flagship-only: on A1, OpenAI's mean answer went 2,691 → 2,166 chars with search on while Anthropic's went 3,473 → 4,479; on the older OpenAI arms retrieval lengthened answers (A2 3,496 → 4,861; A2+ 4,252 → 5,640). Recorded, not explained; at k = 1 it is a shape observation, not a finding.

5.5 The mitigation picture as it stands

What nothing in the corpus mitigates: the source-class split (unmoved by network, effort, or day), the volunteered-material contradictions (both members of each pair confident and unhedged), and Claude Code's 15% retrieval — though the API manipulation converts that last one from a mystery into a configuration question with an owner.


6. The reversals — findings about method

During the study, several early readings were contradicted by later data. They are reported as findings, because they are the strongest evidence in the corpus about what a single run, or a single missing cell, is worth.

  1. The retrieval-policy hypothesis (H2) was restated three times, and the trail is the lesson. First formulation, with three of the four vendor × surface product cells measured: "retrieval policy is a property of the surface, not the vendor" — recorded as confirmed at n = 3 cells, then falsified two days later by the fourth cell, when Codex searched 9/9 (OpenAI's chat and code surfaces are indistinguishable at 100%; Anthropic's differ by 4×). Second formulation: "retrieval rate is set by the vendor × surface cell, and both factors move it" — survives, but only as a description of four numbers. Third formulation, after the API manipulation: the model is not what declines to retrieve; the product layer is. claude-opus-5 on the bare API searched on every one of nine tax questions at 5.4 searches each; the same vendor's coding agent running the same model searched on 4 of 27. Every measured rate survived all three formulations untouched; the causal story attached to them was wrong twice. Three cells produced a confident, well-replicated generalisation; the fourth cell reversed its explanation; the manipulated arm replaced it with a mechanism.
  2. F3 localisation: found, withdrawn, restored as a rate, then broken again. A first pass produced a confident surface difference ("ChatGPT localises the GmbH cue, Claude honours it"); a second pass reversed it (same prompt, same surface, different country — hypothesis withdrawn); a third restored it as a rate (2 of 3 Slovak). Then, six days later, the same prompt went German twice in both network states, and with the reasoning toggle on it went to a third country entirely. The instrument itself was unreliable. Lesson: n = 1 per cell produces confident, publishable, wrong findings, and even an n = 3 rate can fail to reproduce a week later.
  3. "Vendor differences" that were surface differences. The F5 volunteering split (Claude volunteers the depreciation cap, ChatGPT doesn't — the study's "most stable difference" at 13 runs) was labelled a vendor property until Claude Code declined to volunteer it. Same vendor, both sides of the split.
  4. The F8 rebate split read as a reliability difference between surfaces ("Claude reproduces, ChatGPT doesn't") until a same-day control pair showed ChatGPT producing the third figure and Claude, ten minutes later, refusing to commit for the first time in five runs. Both surfaces have produced both figures; the instability is a moving parameter, not a surface trait.
  5. The study's own bookkeeping failed its own standard. An audit of the run ledger found five errors in the records — a corpus total that its own components never summed to, two retrieval tallies never updated, a volunteering count off by one, and a run identifier used by two different runs for seventeen days. None moved a conclusion; all violated the standard that every claim carries its exact n. The cause: the ledger was appended both by hand and by script and had no uniqueness check. Hand-appended records need the same guards as generated ones.
  6. Smaller instances of the same shape: an F9 uncertainty flag seen once on Claude never reproduced; a "retrieval tracks knowledge locus" reading of the search-skip pattern collapsed into "stable in degree, unstable in target"; an interactive picker was first recorded as a version boundary contaminating a comparison, then turned out to be one of several response forms appearing and disappearing at the same setting — variance, not a rollout.

The through-line: between-surface differences that are real reproduce as rates over n ≥ 3; anything asserted from fewer runs, or from an incomplete factor grid, got reversed. That is the operational definition of the D_between-versus-D_within discipline, demonstrated on this study's own claims five times in eighteen days.


7. Limitations

Constraints on what this evidence can support, stated as such:


8. Implications for producers of reference material intended for AI retrieval

Guidance for anyone publishing content — official guidance, professional commentary, structured reference data — meant to be retrieved and cited by AI systems. Each item follows from a finding above; none requires knowing which answer was correct.

  1. Write for the elaboration, not just the answer. Every flat contradiction found in this corpus sits in material the surface volunteered around the question — whether a threshold is gross-or-net of VAT, whether a rebate exists this year, whether a relief was discontinued — while the questions' own traps were corrected 87/87. Reference material that pins down only the headline answer duplicates the one thing surfaces already do consistently; the value is in pinning the surrounding rule cluster a model will volunteer alongside it.
  2. Carry the conditional branches. The F8 fallback branch produced eight values with a 2.7× spread while its headline never moved, across both vendors and three product surfaces. A worked method should enumerate the branch computations — doesn't-qualify paths, partial exemptions, phase-outs — because that is where every surface wobbled.
  3. Date every moving parameter. The most reproduced content failure shape in the corpus is a parameter that changed mid-year (a 40%→50% rebate change, a rebate's existence, a relief's discontinuation, a legislative cutoff date) being served at different vintages by different runs, unflagged — and the API manipulation showed the vintage can flip on the search toggle alone. Content that states the parameter, its effective date, and what it changed from addresses the exact divergence observed.
  4. Never treat an entity name as a jurisdiction signal, and expect no signal at all. F3's "GmbH" resolved to three countries on one surface; F6 showed the no-cue default is the network country on chat and the US on the agentic surfaces and the raw API (0/12 non-US). Tooling that resolves jurisdiction from a question should require or ask for the country — the corpus says asking is what the most stable surfaces do, and they disclose the assumption when they don't ask. It also says questions in which an asker in one country uses another country's instruments are a real, high-variance shape, not an edge case.
  5. The retrieval bottleneck is the product harness, not the model — target it as one. The model underneath every Anthropic surface reaches for search ~100% of the time when the tool is in the call; the same vendor's coding agent ships a configuration in which it reaches 15% of the time on identical questions. The ceiling on how often any retrieval source is consulted is set per-surface by the product layer (system prompts, tool framing, agent defaults), and moving it is a configuration-and-integration question with an owner, not a content-quality problem. Meanwhile the surfaces that retrieve on essentially every question — flagship models at 5.4 and 13.1 searches per question — are where placement is contested several times per question, and where retrieved content is actually read today.
  6. Know which shelf each vendor reads from. In this corpus ChatGPT's citations are predominantly statute and tax authorities; Claude's are predominantly practitioner and commercial commentary. The same content plausibly needs different routes to be read by the two ecosystems. (Candidate, but the most durable observation in the corpus; not coded on the API cells.)

9. Implications for people asking tax questions of AI systems

Each item is traceable to a finding above; none requires knowing which answer was correct.

  1. Name your country — and state — in the question. The one instruction that never failed (16/16 under manipulation). Name the country even when you've named an entity or provision; an entity name alone can silently select a different country's tax system.
  2. Assume nothing about who the answer is for. With no cue, chat at default settings answers for wherever your network appears to be, without saying so (confirmed, one surface); coding agents, raised-effort modes and the raw APIs assume the United States (19 observations across four surface classes and both vendors). If you didn't say where you are, the answer's jurisdiction was chosen for you.
  3. Trust the correction, verify the elaboration. Surfaces reliably fix mis-termed and leading questions. The unreliable material is what they add beyond the question — volunteered thresholds, rebates, reliefs — which is exactly where flat, unhedged contradictions were found.
  4. Ask twice, in fresh sessions, before acting on a figure. The same question produced tax bills differing by a factor of two on one surface within a quarter-hour. Two runs that disagree cost nothing and tell you the figure is contested.
  5. Treat "thorough mode" as sometimes-available, not as the fix. It improves precision and makes assumptions visible, but on one free tier it can consume an entire session allowance on one question, and it may shift the answer away from your own jurisdiction.
  6. A current-year figure, or anything that changed mid-year, is the wobble zone. Stable rules reproduced; moving parameters did not — and whether your answer reflects this year's value can hinge on whether the surface happened to search. Check those against the official source the answer names — and prefer answers that name one.

10. Conclusion

Surfaces are real objects: retrieval rate, source class, jurisdiction policy and disclosure posture reproduce as rates and read as design decisions. The API manipulation measured the largest of them: the surface gap in retrieval (63% → 15% within one vendor) is categorically larger than a full model generation, which moves search intensity (5.4 → 3.3 searches per run) but never the decision to search. The models want to look things up; product layers decide whether they may — which converts the corpus's most alarming cell, a coding agent answering 23 of 27 tax questions from memory uncited, from a model property into a configuration with an owner. Within-surface noise remains large enough that any single answer is unreliable in ways invisible to the person who asked once, and the instability concentrates in exactly four places: moving parameters, volunteered rule statements, conditional branches, and unstated jurisdiction assumptions. The badly-posed question is the solved part; the well-posed question's surroundings are not. Beyond the one manipulated mechanism, everything here is a candidate: formed on the data that suggested it, reported with its exact n and its confounds, and awaiting fresh scenarios, replicated cells and blind coding before it can be more.


Appendix A — the prompt bank, verbatim

All 27 prompts, reproduced exactly as posed. Each family is one underlying question written at three rungs — P1 oblivious, P2 informed layperson, P4 professional — with the family's four design coordinates (territory · knowledge locus · question integrity · tax domain) at its head.

Only the P2 rung has been run. All 262 analysed runs use the nine P2 prompts below; the P1 and P4 prompts are printed for completeness and were not run.

Several prompts contain deliberate errors and mis-stated terminology by design. The Q3 families (F2, F4, F9) use the wrong word for the thing meant, F2 supplies a profit figure where the rule runs on turnover, F5 and F7 lead with premises inviting agreement, and F4 conflates an entity type with a tax election. These are the experimental traps — a "corrected" prompt would be a different experiment.

F1 — Germany · K1 stable rule · Q2 underspecified · personal income

F1-P1

My wife and I both work — we're in Munich. Someone at work said couples can put their tax in together and that it works out better than doing it separately. Is that right? Would it be better for us?

F1-P2

My wife and I are both employed and live in Munich. We earn quite different amounts. Should we file our income tax jointly or separately — which one usually leaves a couple better off?

F1-P4

Married couple, both resident in Munich, both employed, materially different gross salaries. Zusammenveranlagung vs Einzelveranlagung — give me the mechanics and where the crossover sits. Cite the provision.

F2 — Germany · K2 current figure · Q3 mis-termed · VAT

Deliberate traps: "sales tax" for VAT, and a profit figure where the rule runs on turnover.

F2-P1

I've been selling my ceramics at markets around Cologne and last year went better than I expected — I cleared a bit over €25,000 after my costs. At some point do I have to start adding the extra percentage onto my prices? Someone told me there's a limit before you have to.

F2-P2

I'm a small trader in Cologne. What's the sales tax threshold where I have to register and start charging sales tax? I made a little over €25,000 profit last year.

F2-P4

Kleinunternehmer, Cologne. Prior-year figure just over the limit. Confirm the current sales tax registration threshold and the switchover mechanics — and which year's figure governs.

F3 — Germany · K5 indeterminate · Q5 leading · business & entity

Unanswerable without facts only the asker holds; "GmbH" is the only jurisdiction cue.

F3-P1

A friend keeps telling me that if I set up a proper company instead of just working for myself, then any money I leave sitting in the company is barely taxed at all. That sounds too good to be true. Is he right?

F3-P2

I've read that with a GmbH you pay considerably less tax on profits you retain in the company than you would as a sole trader. That's right, isn't it? I'm trying to decide whether to switch.

F3-P4

Client considering GmbH over Einzelunternehmen, intends to retain most profits rather than distribute. My read is that the retained-earnings comparison clearly favours the GmbH. Confirm, and tell me what would flip it.

F4 — United States · K1 stable rule · Q3 mis-termed · business & entity

Deliberate trap: entity type conflated with tax election ("the S-corp rate").

F4-P1

I do handyman work on my own around Sacramento. My sister keeps telling me I should turn myself into a company so I keep more of what I earn. What kind should I set up?

F4-P2

I'm a sole proprietor in California. I want to incorporate as an LLC so I can get the S-corp tax rate. How do I go about that?

F4-P4

Single-member LLC, California, roughly $180k net. Client keeps calling it "the S-corp rate." Give me the election mechanics and the reasonable-compensation exposure. Cite.

F5 — United States · K2 current figure · Q5 leading · personal income

Leading premise that is broadly right but incomplete; needs a figure that moves annually.

F5-P1

I bought a pickup a few months ago for my small landscaping business. A mate told me the whole cost comes off my tax this year — the entire thing. That's how it works, right?

F5-P2

I bought a truck this year for my landscaping business. I can write off 100% of it under the bonus depreciation rule this year, right?

F5-P4

Schedule C landscaper, California. Placed a >6,000 lb GVWR pickup in service this year, roughly 80% business use. My read is that full expensing applies. Confirm the current-year bonus depreciation percentage and the §179 interaction.

F6 — United States* · K5 indeterminate · Q2 underspecified · sales tax

Carries no locale cue at any rung — the pure test of which jurisdiction a surface assumes when nothing points to one. (*Array label; see §4.1.)

F6-P1

I make enamel pins in my kitchen and I've started selling them through a website to whoever orders. Is there some extra tax I'm supposed to be collecting from people when they buy?

F6-P2

I sell handmade goods online. Do I need to be collecting sales tax, and if so where?

F6-P4

Client sells physical goods direct-to-consumer online, first full year, growing quickly. Where do they have to register for sales tax?

F7 — Singapore · K1 stable rule · Q5 leading · GST

Leading premise (overseas client ⇒ zero-rated) that every run so far has declined to accept as stated.

F7-P1

All my customers are abroad — not one of them is here in Singapore. So that extra tax you have to put on invoices doesn't apply to anything I do, does it?

F7-P2

I'm GST-registered in Singapore and all my clients are overseas. Exported services are zero-rated, so I don't charge GST on any of my invoices — that's correct, isn't it?

F7-P4

Singapore GST-registered supplier, all clients offshore, services delivered remotely. My read is blanket zero-rating under the international services provisions. Confirm, and name what would pull any of it back to standard-rated.

F8 — Singapore · K2 current figure · Q2 underspecified · business & entity

Jurisdiction appears only as a currency symbol at P1 and P2.

F8-P1

I'm going properly self-employed and registering a company. If it clears somewhere around S$300,000 in its first year, how much of that do I end up handing over?

F8-P2

I'm incorporating a private limited company. On roughly S$300,000 of profit in its first year, what will it actually pay in corporate tax?

F8-P4

Newly incorporated private limited, first YA, roughly S$300k chargeable income. Current start-up exemption tiers and the effective rate on that.

F9 — Singapore · K5 indeterminate · Q3 mis-termed · personal income

Deliberate trap: "refund" used for reliefs and the assessment outcome.

F9-P1

I moved to Singapore partway through last year for a new job. When it all gets worked out at the end of the year, how much money do I get back?

F9-P2

I relocated to Singapore mid-year for a job. How much of a tax refund can I claim once all the reliefs are applied?

F9-P4

Employee, arrived in Singapore mid-year. Client is asking what their "refund" will be. Give me the residency determination and which reliefs are in play.