Skip to content
LooparaLoopara
Text Tools6 min read1,347 words

OCR: turning a photo of text into text you can edit

The recognition is already on your phone and runs without uploading anything. What makes results come back mangled is nearly always the photograph, and four fixes take five seconds.

Loopara
A phone screen opening a text document from local storage, where recognised text can be pasted and corrected
A phone screen opening a text document from local storage, where recognised text can be pasted and corrected

Short answer

OCR finds the characters in an image and returns them as editable text, and both iPhone and Android run it on the device with nothing installed — tap the text indicator on a photo, or use Google Lens. Poor results are almost always the photograph rather than the software: shoot flat, in even light, with the page filling the frame.

On this page
  1. The one already on your phone
  2. Why does the result come back mangled?
  3. Arabic, and other scripts that behave differently
  4. Turning the text into something you keep
  5. A five-step routine that works every time
  6. When is a website the right choice?
  7. The short version

A photographed page, a screenshot of a slide, a receipt you need the numbers from. OCR is optical character recognition — software that finds the letters in an image and returns them as text you can select, search and edit — and on a modern phone it is already installed, which is the part most people miss.

Almost nobody needs a website for this any more. The recognition runs on the device, in about a second, on hardware built for it.

The one already on your phone

Before installing anything, try what is there. Both platforms ship OCR into the camera and the photo library.

On iPhone, open a photo, and if there is text in it a small indicator appears in the corner. Tap it and the text becomes selectable, exactly like text on a web page — copy it, search it, or tap a phone number to call it. It works in the Camera app before you take the picture, too.

On Android the equivalent lives in Google Lens, reachable from the photo viewer or the camera, with the same select-and-copy behaviour.

Three things are worth knowing about the built-in version:

  • It runs on the device. The image is not uploaded, which matters for a passport, a payslip or a contract.
  • It handles multiple scripts, including Arabic, though accuracy varies more than it does for Latin text.
  • It is not a document scanner. You get the text, not a searchable PDF with the layout preserved.

That third point is the real dividing line between the built-in tool and everything else.

What you needUse
A few lines out of a photoBuilt-in, on the phone
A whole page as editable textBuilt-in, then paste into an editor
A searchable PDF that keeps the layoutA scanning app that writes a text layer
Fifty pages in one goA desktop tool with batch OCR
A table with its structure intactSpecialist tooling, and expect to fix it
OCR gives you characters. Keeping the layout, the reading order and the table structure is a separate and much harder problem.

Why does the result come back mangled?

Almost always the image, not the software. Four causes cover nearly everything.

The photograph is at an angle. OCR expects roughly horizontal lines of text. A page shot from 30 degrees off produces a recognisable image and poor recognition. Straighten before you scan — every phone's document mode does this automatically, which is reason enough to use it.

The lighting is uneven. A shadow across half the page, or glare from a window, removes the contrast between ink and paper that the recogniser depends on. Diffuse light and no flash beats a bright direct light.

The resolution is too low. A screenshot of a screenshot, or a photo taken from too far away. As a rough guide, a character needs to be around 20 pixels tall to be read reliably; below that, accuracy falls off sharply.

The text is not typeset. Handwriting, decorative fonts and heavily stylised logos are a different problem from printed text, and general OCR handles them badly. Handwriting recognition exists and is separate.

The fix for the first three is the same and takes five seconds: use the document mode, on a flat surface, with the page filling the frame.

Arabic, and other scripts that behave differently

Worth its own section, because the advice for Latin text does not transfer cleanly.

Arabic is cursive, letters change shape by position, and diacritics are optional marks that recognisers may add, drop or misread. Accuracy on clean printed Arabic is good and falls faster than Latin does as image quality drops. Expect to proofread, particularly numbers and names.

Right-to-left text also creates a reading-order problem that pure character recognition does not solve. A page mixing Arabic prose with Latin product names or figures can come back with the runs in the wrong order, even when every character is correct. That is not a recognition failure; it is a layout failure, and it needs a human eye.

The practical approach for mixed-script documents is to work in small pieces — a paragraph at a time rather than a page — because a short run is easy to check and a long one hides errors in the middle.

Turning the text into something you keep

Recognised text is not yet a document. Two steps make it useful.

  1. Paste it into a plain-text or Markdown editor first, not a word processor. You want to see the raw result, including the line breaks OCR inserted at the end of every visual line rather than every paragraph.
  2. Join the lines. This is the single most common cleanup, and it is a find-and-replace: a line break not followed by a blank line usually belongs to the previous sentence.

Working in Markdown is convenient here because the file stays readable and portable while you fix it — an app like Mdora opens the file directly from local storage, so the corrected text is a file you own rather than a document trapped in an app. Our guide to cleaning up messy text covers the other operations that follow, and file tools covers what to do with the result.

If the goal is to listen rather than read, the recognised text feeds straight into text-to-speech — that path is covered under guides.

A five-step routine that works every time

Once the setup is right the job takes under a minute, and the order matters more than the tool.

  1. Flatten the page and put it on a surface with even light. No flash — it creates glare exactly where the text is.
  2. Use document mode rather than a plain photo. It straightens the perspective and crops to the page, which fixes 2 of the 4 common causes of bad results.
  3. Fill the frame. Text should be large in the image; a page shot from a metre away loses the detail the recogniser needs.
  4. Check 3 places before trusting it — the first line, a line in the middle, and any numbers. Errors cluster in low-contrast areas rather than spreading evenly.
  5. Fix the line breaks before you do anything else with the text.

Step 4 is worth the 15 seconds. OCR fails confidently: it does not flag uncertainty, it simply returns a plausible wrong character, and a misread digit in an account number looks exactly like a correct one.

When is a website the right choice?

Rarely, and the exceptions are specific.

A batch of scanned pages, an unusual language your phone does not support, or a PDF that needs a text layer adding while keeping its pages intact — those are genuine reasons to use a desktop tool or a service.

What is not a good reason is convenience for a single photo, because uploading is a real cost. An identity document, a bank statement or a signed contract sent to an unknown processor is a decision with consequences, made to save the ten seconds that opening the Photos app would have taken.

The rule that covers almost every case: if the document is one you would not email to a stranger, do the OCR on your device. The tool is already there, it is fast, and the image never leaves.

Apple documents the underlying framework in Vision, which is the same recognition the Photos app uses — worth a look if you want to know what is actually running.

The short version

Try the built-in OCR first; it is on the phone, it runs locally and it is good. Shoot flat, in even light, filling the frame, using document mode. Expect to join the line breaks afterwards, and to proofread Arabic more carefully than English.

Reach for a website only for batches, unusual languages, or a searchable PDF — and never for a document you would not hand to a stranger.

Frequently asked questions

Does phone OCR upload my photo anywhere?
The built-in recognition on both platforms runs on the device, so the image itself is not sent. That is the main reason to prefer it over a website for anything sensitive, such as an identity document or a bank statement.
Why does the text have a line break at the end of every line?
Because OCR reports what it sees, and it sees visual lines rather than paragraphs. Joining them is the standard cleanup — a line break not followed by a blank line usually belongs to the sentence before it.
How well does OCR handle Arabic?
Clean printed Arabic recognises well, but accuracy falls faster than Latin as image quality drops, and diacritics may be added or lost. Mixed Arabic and Latin can also come back in the wrong reading order even when every character is right.
Can OCR give me a searchable PDF rather than plain text?
Not the built-in tool, which returns text without layout. A searchable PDF needs a scanning app or desktop tool that writes an invisible text layer over the original page image.

Sources

  1. Recognizing text in images — Vision frameworkApple Developer
  2. Search what you see with Google LensGoogle Photos Help
  3. Use Live Text with your iPhone cameraApple Support

Published by

Loopara

Practical guides, free tools, workflows, and resources for productivity, files, images, video, text, creators, and everyday digital tasks.

About the publication

Related reading

Keep going