Skip to content
GPU-assistedUpload required

Remove Text from an Image

There are two kinds of text in a picture and only one of them should go. A caption laid over the top is an addition to the photograph. A shop sign, a book cover, a road marking, or a number plate is part of what was photographed. To you the difference is obvious. To a model reading a flat grid of pixels, both are simply lettering, and telling them apart is a judgment it makes from context rather than a distinction it can see.

Use this without the search next time. Prathom Workbench puts Prathom's tools in your toolbar.

Add to Chrome — free

Drop an image here, or click to browse

Up to 25 MB. Uploaded temporarily, deleted after processing.

GPU-assisted — results vary slightly between runs.

What it does

  • Overlaid captions and lettering removed
  • Surface beneath the text rebuilt
  • Scene signage explicitly protected in the instruction
  • Roughly one megapixel output

How to use Remove Text

  1. 1

    Note every piece of text in the frame first

    Include the ones you want kept. Those are the ones at risk, and knowing where they are is what makes checking the result quick.

  2. 2

    Run the removal

    The image is read at about one megapixel and repainted with overlaid text gone and the surface beneath it rebuilt.

  3. 3

    Read what is left

    Not glance at it, read it. Partial removal and text replaced by text-shaped marks are the two failures here, and both survive a glance.

How it works

Your image is scaled so its total area is around one megapixel, keeping the original proportions. An image-editing model reads that copy along with a fixed instruction: remove the overlaid text, captions, and lettering and rebuild the surface underneath so the photograph beneath reads as continuous, while preserving the subject, the objects, the signage that belongs to the scene, the lighting, and the composition, and writing nothing new.

The negative instruction names remaining letters, garbled lettering, smudged patches, altered subjects, and removed signage.

The repaint runs at sixty five per cent strength — lower than object or people removal, because text is thin and covers very little, so less has to be rebuilt. The seed is not fixed, so a second run reconstructs differently.

Overlay and content look the same

This is the distinction the whole page turns on, and it is genuinely hard rather than a shortcoming of this particular model.

When you look at a picture with a caption across it, you are not reading pixels. You are using knowledge about how images are made: captions are horizontal, sit in a sans-serif face, ignore the perspective of the scene, sit at a uniform brightness, and are placed in the flat areas. A shop sign obeys the perspective of the wall it is on, carries the same light as everything around it, and is usually smaller and less tidy.

A model can use those cues too, and it does. That is why the instruction works at all, and why captions on skies are removed cleanly while a street name plate usually survives.

But they are cues, not categories. A sign photographed straight on, in even light, in a clean typeface, looks exactly like an overlay. A caption set at an angle over a photograph looks like part of it. In those cases the model guesses, and it has no way to ask you which you meant.

So the working rule is simple: any text in the frame that must survive is text you have to check, every time.

The easiest surface on the site

Text removal has one large advantage over the other cleanup jobs here.

Captions are placed where they can be read, which means over the calmest part of the picture: a sky, an out-of-focus background, a wall, a dark area, a deliberately blurred band. Designers put them there on purpose.

Those are precisely the surfaces that are cheapest to invent. A gradient sky has no structure that can be got wrong. A blurred background has no detail to match. So the reconstruction usually succeeds completely, and this is the removal most likely to work on the first attempt.

The exception is text placed across a detailed scene, which is common in memes and screenshots. Now the model is rebuilding real content through the strokes of each letter, and the result is judged by the same standards as any other object removal.

Read the result rather than looking at it

Both failure modes here survive a glance, which makes checking unusually important.

The first is partial removal. Most of the caption goes and one character stays, often at an edge or where it crossed something high in contrast. At a glance the text is gone; at reading distance there is a stray letter floating in a sky.

The second is substitution. The model repaints the region as something text-shaped — marks with the spacing, weight, and line rhythm of writing that resolve into nothing when read. This happens because a caption-shaped area usually does contain writing, and the model is completing the pattern.

Both are obvious the moment you actually try to read what is there, and both are invisible if you only check that the picture looks clean.

Publication gate

This page ships once the workflow has been run against a caption on a sky, a caption across a detailed photograph, a scene containing a legible shop sign, and an image with text at the frame edge, with every output read at full size rather than inspected.

Examples

Caption over a plain area

quote.jpg - 1600x1600, white text across a sky
prathom-remove-text.png - 1024x1024, sky restored

The easiest removal on the site. A gradient sky has no structure to reconstruct, so the surface underneath is nearly free to invent.

Caption over a detailed scene

meme.jpg - 1200x900, lettering across a busy photograph
clean-scene.png - 1160x870

Harder, and the case where in-scene text is most at risk, because the model is already looking for lettering across the whole frame.

Frequently asked questions

Will it remove the signs and labels inside my photograph too?

Sometimes, and this is the failure to check for. The instruction asks for overlaid text specifically and names scene signage as something to preserve, but overlay and content are not visually distinct — both are letterforms on a surface. Anything in the frame that should stay readable needs verifying afterwards.

Why are there a few letters left behind?

Because removal happens through repainting rather than detection, so nothing tracks whether every character has gone. A letter with unusual contrast, one overlapping a busy region, or one at the edge of the frame can survive while its neighbors do not. The negative prompt names remaining letters, which reduces it rather than preventing it.

Can it remove a watermark?

Technically it may, and that is a separate question from whether you should. Removing a mark from an image you do not have a license for does not give you one, and the mark's absence is not a defense. On your own images, or your own watermark, this is an ordinary cleanup job.

Why does the text sometimes come back as gibberish?

Because the model has learned that regions shaped like captions contain writing, so when it repaints one it can produce something with the rhythm and weight of text rather than the surface underneath. It looks like letters until you try to read it. Rerunning usually resolves it, since each run reconstructs differently.

Is the AI Remove Text tool free, and do I need an account?

No sign-up. The remove text workflow takes an image up to 25 MB, runs it on the GPU host and hands back a file you can download directly. Nothing is charged and nothing is stamped onto the output. Because no account exists, nothing keeps your history either, so save the result within the thirty minutes the job is retained.