JPG to Word Converter — Image to Word

Two different jobs, and this page does both without pretending they are the same one: put the picture into a real Word document, or read the words out of it as text you can edit. Nothing is uploaded either way.

Image to Word converter

Convert

What do you want in the document?

Embed mode makes no network request at all. The text in the picture will not be selectable or searchable in the finished document — if you need that, switch to the other mode.

Supports: JPGJPEGPNGWEBP
+2 more GIFBMP
Document layout

Nothing is uploaded. This tool runs entirely in your browser, so the file never leaves your device. There is no server to send it to, no copy kept anywhere, and it keeps working if you lose your connection after the page has loaded.

Accepts .jpg .jpeg .png .webp .gif .bmp You get: A .docx file — the picture, or its text Free, no sign-up

Two different jobs wearing one name

Search for "JPG to Word" and you will find two groups of people asking for opposite things. One has a picture — a photo, a diagram, a screenshot, a signed page — and needs it inside a document they can print or send. The other has a picture of writing and needs the writing back as text they can edit. Most converters pick one silently and let the other group find out after downloading.

This page asks. Embed puts the image into the document. OCR reads the words out of it. They produce genuinely different files, so the mode you choose changes the button, the options and the description above it — and neither is described as the better one.

Embed mode: the picture, in a real document

Each image is written into a genuine .docx — an OOXML package with a document part, a styles part, a relationship for every picture and the image bytes themselves stored inside. Word opens it, prints it and lets you move or resize the picture like any other. It is not a renamed image and not a PDF in disguise.

Pictures are scaled down to fit the page's text area and centred, keeping their proportions. They are never scaled up: a small logo stretched to page width looks worse than it did to start with, and nobody asked for that. The original bytes go in unchanged, so a JPEG is not re-compressed on the way and nothing is lost.

The thing this mode cannot do is the thing its name suggests to half the people who search for it: the writing in the picture stays part of the picture. It cannot be selected, searched, spell-checked or edited. That is a property of putting an image in a document, not a shortcoming of this page, and it is why the other mode exists.

OCR mode: the words, as editable text

OCR — optical character recognition — looks at the shapes in the image and works out which characters they are. The result is real text: paragraphs you can click into, edit, search and spell-check, written into the document as ordinary Word paragraphs rather than as a picture.

The engine is Tesseract, compiled to WebAssembly, together with a trained data file for the language you choose. Both are downloaded the first time you pick OCR and cached by your browser afterwards. Recognition then runs inside the page, on your own device. Your image is not part of that download and is never sent anywhere.

Because recognition is never perfect, the text lands in an editable box rather than going straight into a file. Each image gets a confidence figure, and the paragraphs read least reliably are flagged. The document is built from whatever is in those boxes when you press Download, so a correction you make is a correction in the file.

What OCR will not do

It will not reproduce your layout. The words come back in reading order as plain paragraphs: columns are flattened into one, tables become lines of text, and fonts, sizes, colours, borders and any images in the original are gone. If you need the page to look the same, Embed mode is closer to what you want, because it keeps the appearance and gives up the editing.

Accuracy depends almost entirely on the picture. Straight, evenly lit, reasonably large printed text reads very well. A photo taken at an angle, faint or low-contrast printing, decorative fonts, heavy scanning noise and anything handwritten all degrade it — and the characteristic failure is not a gap but a confidently wrong word. That is exactly why the confidence figures are shown and why the text is editable before it becomes a document.

Choosing the wrong language makes it much worse rather than slightly worse, because the model is deciding which characters can exist at all. Twenty-four languages are offered; pick the one the document is actually written in.

Common mistakes

Choosing OCR mode for a photo that is not of writing — a landscape, a product shot, a face — is the most common mismatch. Recognition can still return something: Tesseract scores confidence in the shape it matched, not in whether writing exists at all, so a picture with no text can occasionally produce a stray character at a high confidence score. That is exactly why this page checks the volume of recognised text as well as the confidence figure, and treats a result under twelve characters as nothing found rather than presenting a one-letter "document" as real text. If the message says almost nothing was found, Embed mode is very likely the one you actually want.

Picking the wrong OCR language is the second common mistake, and its effect is easy to underestimate: the model is not just less accurate in the wrong language, it is actively guessing among the wrong set of characters, so results can come back confidently garbled rather than obviously wrong. If a document mixes two languages, recognising it twice — once per language, keeping the better result for each part — works better than picking one and hoping.

Edge cases

An image with genuinely no readable text — a blank page, a photo of a scene — is reported as having almost nothing found rather than turned into a document with a stray guessed character in it, for the reason described above. An image where OCR fails outright, for instance because it is corrupted in a way the recognition engine cannot open, is left out of the document entirely and marked as could not be read, rather than silently skipped. Switching between Embed and OCR clears any recognised text, because text read for one image set or one language should not survive into a different one and be built into a document by mistake.

If you only need the picture converted, not placed inside a document, JFIF to JPG, JFIF to PNG and WebP to JPG handle the image formats directly. For a PDF instead of a Word document, see JFIF to PDF, which shares this page's no-upload approach but writes PDF pages rather than a .docx. Neither of those does OCR — that stays specific to this page.

What is downloaded, and what is not sent

Embed mode makes no network request whatsoever. Load the page, disconnect, and it still builds documents.

OCR mode fetches the engine and the language data from a public CDN. That request tells the CDN your IP address, as any request for a file does — it does not tell it anything about your image, which is read by code running in your own browser and never uploaded. This page has no upload endpoint to send it to.

Formats and limits

JPG, JPEG, PNG, WebP, GIF and BMP, up to twenty images at 25 MB each. The format is decided by reading the first bytes of the file, not by trusting its extension or the type the browser reports — both of those are attacker-influenced the moment someone is talked into converting a file they were sent. Anything that fails that check is rejected by name, with the reason, rather than producing a broken document.

Everything is assembled in memory in your browser, so the practical ceiling is your device rather than a server quota. A phone will reach it well before a laptop, and OCR in particular is real work: expect a second or two per clean page and longer for large or difficult ones.

Frequently asked questions

Which of the two modes do I want?

If you want to put a picture into a document — a photo, a diagram, a screenshot, a signed page — use Embed. If the picture is of writing and you want that writing back as text you can edit, use OCR. The quick test: if you would be annoyed that you cannot select the words in the result, you want OCR.

Is the result a real Word file?

Yes, in both modes. It is a genuine .docx — an OOXML package with a document part, a styles part and, in Embed mode, the image stored inside it. Word, LibreOffice, Google Docs and Pages all open it and let you edit it normally. Nothing here renames a file and hopes.

Can I edit the text after using Embed mode?

No, and that is the honest limit of that mode. Embed puts the picture into the document; the writing in it is part of the picture, exactly as it is in the original file. You can move the image, resize it or type around it, but you cannot click into the words. If you need the words, use OCR mode.

How accurate is the text recognition?

On a clean, straight, reasonably large scan of printed text it is usually very good. On a photograph taken at an angle, on faint or low-contrast text, on unusual fonts and on handwriting it degrades, sometimes badly — and it degrades by inventing plausible words rather than leaving gaps. The page shows a confidence figure for each image and flags the paragraphs it is least sure about, so you can see where to look.

Will it keep my layout, columns, tables and colours?

No. OCR mode returns the words in reading order as ordinary paragraphs. Columns are flattened, tables come out as lines of text, and fonts, sizes, colours and images are not reproduced. Anything that promises a layout-faithful Word file from a photograph is either using a full document-reconstruction service or overstating what it does.

Are my images uploaded anywhere?

No. Both modes run on your own device. Embed mode makes no network request at all. OCR mode downloads the recognition engine and the language data from a public CDN, and the recognition itself then runs inside your browser — your image is never sent anywhere, in either mode. This page has no upload endpoint.

Why does OCR need a download at all?

Because recognising text needs a trained model, and the model is the download. The engine is Tesseract, compiled to WebAssembly, plus a data file for the language you pick. It is fetched the first time you choose OCR and cached by your browser afterwards. If you only ever use Embed mode, none of it is fetched.

Which languages can it read?

Twenty-four are offered, covering the Latin, Cyrillic, Arabic, Devanagari, CJK, Thai and Tamil scripts among others. Pick the language the document is actually written in — recognition quality drops sharply when the wrong one is chosen, because the model is guessing at which letters exist. Chinese, Japanese and Korean packs are considerably larger downloads and the list says which those are.

Can I correct mistakes before I download?

Yes, and you are meant to. The recognised text appears in an editable box for each image, with a confidence figure and a warning on the paragraphs read least reliably. The document is built from whatever is in those boxes when you press Download, so a fix you make there is a fix in the file.

Which formats and how many images?

JPG, JPEG, PNG, WebP, GIF and BMP, up to 20 images at 25 MB each. The format is checked by reading the first bytes of the file rather than trusting its name. Everything is assembled in your browser, so a phone will run out of memory long before a laptop does.

About this converter

The document is built to the OOXML specification and its structure is checked automatically — every required part present, every relationship resolving, every picture declared. It has not been opened in Microsoft Word itself, because that is not available here; if a file ever fails to open for you, that is a real defect and worth reporting. Recognised text is produced by Tesseract, an open-source engine that runs in your browser, and its output is a best effort rather than a transcript: do not rely on it for anything legal, medical or financial without reading it against the original first.

JPG to PPT

Images into a real PowerPoint deck, one picture per slide.

Document Tools

JFIF to PDF

JFIF images into a PDF, one per image or all in one.

Document Tools

PPT to Word

Slide text and notes into a .docx, in slide order.

Document Tools

HEIC to JPG

iPhone HEIC photos into ordinary JPGs, without uploading them.

Image Tools

JFIF to JPG

Turn .jfif downloads into ordinary .jpg files, in your browser.

Image Tools

PPT to HTML

Turn a deck into clean, readable HTML sections.

Document Tools

More in Document Tools All tools