Foundations

What these systems are structurally bad at

Shared methods · A shared method; linked tool guides explain the exact steps.

Five limits that follow from how the thing is built, so they do not go away with a better prompt or a newer model. Plan around them.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Plenty of things these systems handle badly today will be handled well in two years. This page is about the other category: limits that fall out of the architecture, which no amount of prompting technique moves.

Hold onto that distinction. Effort spent working around a structural limit is well spent; effort spent trying to prompt one away is not.

It cannot tell you how sure it is

A model produces a distribution over next tokens and samples from it. Nothing in that process yields a calibrated statement about whether the resulting claim is true.

So a well-evidenced fact and a plausible guess arrive identically: same fluency, same confidence, same length. When you ask how certain it is, you get a description of certainty generated by the same process that generated the claim, which is not a measurement of anything.

This is the limit that makes every other one dangerous. A system that failed loudly would be far easier to work with.

It has no state between requests

Nothing persists. Every turn re-sends the whole conversation, and features that seem to remember — memory, project knowledge — are separate machinery that writes text down and pastes it back in.

Practically: it cannot learn from your correction in any lasting sense. Explain your preference in ten conversations and you will explain it in the eleventh, unless something wrote it into a store.

Attention thins as context grows

Anthropic describes this as context rot: "as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases."

Three architectural reasons, all in their engineering write-up. Transformers create n² pairwise relationships for n tokens, so longer contexts stretch the model's ability to hold those relationships. Models carry what Anthropic calls an "attention budget". And training data contains shorter sequences than the maximum context window, leaving "fewer specialized parameters for context-wide dependencies".

A million-token window is a capacity figure, and capacity is a different claim from uniform quality across the whole span.

It is not a function

Identical input does not guarantee identical output. The API reference is explicit even about the setting that used to make people believe otherwise: "Note that even with temperature of 0.0, the results will not be fully deterministic."

Anything requiring reproducibility — audit trails, regression tests that assert exact text, a process that must give the same answer twice — needs designing around this. See why answers vary.

Its view of the world is skewed and unauditable

The APA's advisory puts both halves plainly: "The foundational training data for most large language models (LLMs) are not publicly available, which precludes their systematic evaluation." And those data are "not globally representative, consisting primarily of English-language and Western-centric content from the internet."

Two consequences, and they are separate. The skew means coverage is uneven in ways that track what gets written on the internet in English, so a topic under-represented there gets thinner treatment with no signal that it has. The unauditability means neither you nor anyone outside the lab can check where the gaps fall.

What follows

These limits are stable enough to design against.

Keep context small and deliberate, because quality tracks signal density more than volume. Externalise anything that must persist. Verify against something outside the model, especially where you lack the expertise to review. Prefer tasks where checking the answer is cheaper than producing it — that asymmetry is where these systems genuinely pay, and it is why they suit drafting, summarising and generating candidates better than final judgements.

Try this

Ask something you know is barely documented in English. Watch what comes back: fluent, structured, appropriately hedged, and thinner than it appears. Compare against a question in a well-covered area. The difference in tone between the two is close to nothing, which is the entire lesson.

What goes wrong

Prompting harder at a structural limit. Elaborate instructions aimed at making it "only say things it is sure of" cannot succeed, because the underlying signal is absent.

Reading a bigger context window as a bigger reliable window. Capacity grew. Attention behaviour did not change shape.

Assuming the next model fixes it. Some of these have improved and will improve further. Statelessness, non-determinism and absent self-knowledge follow from how the system works.

Treating uneven coverage as even. The failure is silent by construction: sparse training data produces confident output, not an error message.

How to check it worked

Take a task you rely on and name which of the five it is exposed to. If you cannot name one, either the task is genuinely well-suited — checkable output, short context, well-covered subject — or you have not yet found the exposure. Both are worth knowing, and the second is the one that shows up later on its own schedule.

Sources

  1. Effective context engineering for AI agents — Anthropic Tier 1 2026-09-02
  2. Create a Message — Claude API reference Tier 1 2026-09-02
  3. APA Health Advisory on the Use of Generative AI Chatbots and Wellness Applications for Mental Health (November 2025) Tier 1 2026-09-02