← All posts
5 min read

How to Chat With Images and Scanned Documents Using AI: A Step-by-Step Guide

Chatting with images and scanned documents means asking a photo or a scan questions instead of retyping or squinting at it. Here's how it works and how to get answers you can trust.

What does it mean to chat with images and scanned documents?

To chat with images and scanned documents, a tool first looks at the picture — a photo of a whiteboard, a scanned contract saved as a JPG, a screenshot of a chart — and produces a caption or transcription of what's in it. From there you can ask about it in plain language: “What does this scanned page say about the deadline?” or “Summarize the diagram in this photo,” instead of retyping the text yourself or zooming in on a blurry image looking for one detail.

Under the hood it works the same way as chatting with any other document: once the image has been described in text, that text is split into chunks, each chunk is embedded, and your question retrieves the most relevant ones before a language model writes an answer from them. We cover the general mechanism in our RAG explainer; for an image the only difference is the extra captioning step that turns a picture into searchable text in the first place.

Step by step: chat with an image or scanned document

  1. Pick a tool that accepts standalone images (JPG, PNG) as a source, not just PDFs and text files.
  2. Upload the photo or scan and wait a few seconds while it's captioned and indexed.
  3. Ask a broad question first, like “What is shown in this image?”, to confirm it read the picture correctly.
  4. Follow up with specific questions — a figure from a scanned table, a label on a diagram, a line of text from a photographed page.
  5. Add related files to the same project so a single question can draw on the image and your other sources together.

What AI actually sees in a photo or scan

An image-captioning model reads a picture the way a person skimming it would: it transcribes printed or typed text reasonably well, describes charts and diagrams, and picks up labels and headings. It's less reliable on messy handwriting, and a dark, blurry, or heavily angled photo gives it less to work with than a flat, well-lit one. The practical takeaway is simple — a clear photo taken straight-on, or a proper scan instead of a hasty phone snap, turns into a far more accurate answer than a rushed shot of a page.

Combine photos of notes with your other study material

This is exactly where the workflow pays off for students: studying with AI from your own notes already means uploading photos of handwritten pages alongside lecture slides, and having both readable in the same project means a question can pull from a photographed page and a typed PDF in one answer. The same captioning step also covers images embedded inside a slide deck, as described in how to chat with Office files — a chart on a slide becomes searchable text too, not just the bullet points around it. Once everything is indexed, generating a quiz from the mixed set tests recall across photos and text together.

Try it with your own photo or scan

Doxylets you add a photo or scanned document the same way you'd chat with any other document: upload it as a source, wait for it to be captioned and indexed, and start asking. It's part of the same AI document chat and document Q&A workflow, so a photographed page sits alongside your PDFs, slides, and web links in one project, with citations back to the source you actually uploaded.

Start chatting with your documents

Upload a file or paste a link and ask your first question in seconds. Free to start.

Get started free

We use cookies for analytics to understand how Doxy is used and improve it. You can accept or decline — declining still lets you use Doxy normally. See our Privacy Policy.