Files, images and PDFs: what actually gets read
Claude
This page covers tools outside your selection. You can still read it. Find matching guides
A PDF over 100 pages has its text read and its charts ignored, silently. Knowing what is extracted from what saves you trusting an answer built on half a document.
Attaching a file feels like handing over the document. What actually happens depends on the format and, for PDFs, on the page count — and one of those thresholds changes the answer you get without changing anything you can see.
The 100-page cliff
For a PDF of 100 pages or fewer, Claude analyses the text and the visual elements: images, charts and graphics.
From 101 pages up to the 1000-page ceiling, it processes text only and does not analyse visual elements.
Nothing announces the crossing. Ask about a chart on page 140 of a 200-page report and you will get an answer, because the surrounding text mentions the chart. The answer will be built from the caption and the prose around it rather than from the figure. If the figure is where the information actually lives, the answer is reconstruction.
What each format gives up
PDFs get the fullest treatment, subject to the page rule above.
Everything else — DOCX, ODT, RTF, EPUB, HTML, CSV, TXT, JSON, XLSX — is text extraction only. Embedded images are not processed. A Word document whose argument rests on a diagram arrives as the words around the diagram.
Images in JPEG, PNG, GIF and WebP are read directly. Anthropic recommends 1000×1000 pixels or larger, and accepts up to 8000×8000.
The asymmetry catches people out: exporting that DOCX to PDF before uploading genuinely changes what the model can see, as long as the result stays under 100 pages.
Limits, so you know which wall you hit
In a chat, files can be up to 500 MB each, with a maximum of 20 files per chat. PDFs cap at 1000 pages regardless of size.
Files uploaded to a project's knowledge base are capped lower, at 30 MB each, though the number of files is limited only by what the context window holds.
Try this
Take a PDF you have already asked questions about. Ask: "How many pages is this document, and did you analyse its images and charts, or only its text?" Then check the page count yourself. If it is over 100, every answer you have had about a figure in it was inferred from surrounding text.
What goes wrong
Assuming a long PDF was fully read. This is the failure that produces confident wrong numbers, because a figure's caption often states a different thing from the figure.
Uploading a spreadsheet and expecting spreadsheet behaviour. Text extraction gives the model the cell contents, not the formulas or the relationships between sheets. For anything where the calculation matters, say what the calculation is.
Attaching everything at the start. Every attached file stays in the context window for the rest of the conversation, costing the same on turn forty as on turn one. Attach the one document you are working on and leave the folder it lives in alone.
Treating a scan as a document. A scanned page is an image of text. Under 100 pages it is analysed visually and usually works; above that threshold, a scanned PDF has no extractable text layer to fall back on.
How to check it worked
Ask it to quote the sentence or read the figure label it used, and compare against the document. A quote you can find is evidence the material was reached. A fluent summary that matches the document's general subject is not, and on a long PDF that distinction is the difference between a fact and a plausible reconstruction.
Sources
- Upload files to Claude — Anthropic Help Center Tier 1 2026-09-02
- PDF support — Claude Platform Docs Tier 1 2026-09-02
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.