Two different jobs wearing one name
Search for "JPG to Word" and you will find two groups of people asking for opposite things. One has a picture — a photo, a diagram, a screenshot, a signed page — and needs it inside a document they can print or send. The other has a picture of writing and needs the writing back as text they can edit. Most converters pick one silently and let the other group find out after downloading.
This page asks. Embed puts the image into the document. OCR reads the words out of it. They produce genuinely different files, so the mode you choose changes the button, the options and the description above it — and neither is described as the better one.
Embed mode: the picture, in a real document
Each image is written into a genuine .docx — an OOXML
package with a document part, a styles part, a relationship for
every picture and the image bytes themselves stored inside. Word
opens it, prints it and lets you move or resize the picture like
any other. It is not a renamed image and not a PDF in disguise.
Pictures are scaled down to fit the page's text area and centred, keeping their proportions. They are never scaled up: a small logo stretched to page width looks worse than it did to start with, and nobody asked for that. The original bytes go in unchanged, so a JPEG is not re-compressed on the way and nothing is lost.
The thing this mode cannot do is the thing its name suggests to half the people who search for it: the writing in the picture stays part of the picture. It cannot be selected, searched, spell-checked or edited. That is a property of putting an image in a document, not a shortcoming of this page, and it is why the other mode exists.
OCR mode: the words, as editable text
OCR — optical character recognition — looks at the shapes in the image and works out which characters they are. The result is real text: paragraphs you can click into, edit, search and spell-check, written into the document as ordinary Word paragraphs rather than as a picture.
The engine is Tesseract, compiled to WebAssembly, together with a trained data file for the language you choose. Both are downloaded the first time you pick OCR and cached by your browser afterwards. Recognition then runs inside the page, on your own device. Your image is not part of that download and is never sent anywhere.
Because recognition is never perfect, the text lands in an editable box rather than going straight into a file. Each image gets a confidence figure, and the paragraphs read least reliably are flagged. The document is built from whatever is in those boxes when you press Download, so a correction you make is a correction in the file.
What OCR will not do
It will not reproduce your layout. The words come back in reading order as plain paragraphs: columns are flattened into one, tables become lines of text, and fonts, sizes, colours, borders and any images in the original are gone. If you need the page to look the same, Embed mode is closer to what you want, because it keeps the appearance and gives up the editing.
Accuracy depends almost entirely on the picture. Straight, evenly lit, reasonably large printed text reads very well. A photo taken at an angle, faint or low-contrast printing, decorative fonts, heavy scanning noise and anything handwritten all degrade it — and the characteristic failure is not a gap but a confidently wrong word. That is exactly why the confidence figures are shown and why the text is editable before it becomes a document.
Choosing the wrong language makes it much worse rather than slightly worse, because the model is deciding which characters can exist at all. Twenty-four languages are offered; pick the one the document is actually written in.
Common mistakes
Choosing OCR mode for a photo that is not of writing — a landscape, a product shot, a face — is the most common mismatch. Recognition can still return something: Tesseract scores confidence in the shape it matched, not in whether writing exists at all, so a picture with no text can occasionally produce a stray character at a high confidence score. That is exactly why this page checks the volume of recognised text as well as the confidence figure, and treats a result under twelve characters as nothing found rather than presenting a one-letter "document" as real text. If the message says almost nothing was found, Embed mode is very likely the one you actually want.
Picking the wrong OCR language is the second common mistake, and its effect is easy to underestimate: the model is not just less accurate in the wrong language, it is actively guessing among the wrong set of characters, so results can come back confidently garbled rather than obviously wrong. If a document mixes two languages, recognising it twice — once per language, keeping the better result for each part — works better than picking one and hoping.
Edge cases
An image with genuinely no readable text — a blank page, a photo of a scene — is reported as having almost nothing found rather than turned into a document with a stray guessed character in it, for the reason described above. An image where OCR fails outright, for instance because it is corrupted in a way the recognition engine cannot open, is left out of the document entirely and marked as could not be read, rather than silently skipped. Switching between Embed and OCR clears any recognised text, because text read for one image set or one language should not survive into a different one and be built into a document by mistake.
Related converters
If you only need the picture converted, not placed inside a document,
JFIF to JPG,
JFIF to PNG and
WebP to JPG handle
the image formats directly. For a PDF instead of a Word document, see
JFIF to PDF, which
shares this page's no-upload approach but writes PDF pages rather
than a .docx. Neither of those does OCR — that stays
specific to this page.
What is downloaded, and what is not sent
Embed mode makes no network request whatsoever. Load the page, disconnect, and it still builds documents.
OCR mode fetches the engine and the language data from a public CDN. That request tells the CDN your IP address, as any request for a file does — it does not tell it anything about your image, which is read by code running in your own browser and never uploaded. This page has no upload endpoint to send it to.
Formats and limits
JPG, JPEG, PNG, WebP, GIF and BMP, up to twenty images at 25 MB each. The format is decided by reading the first bytes of the file, not by trusting its extension or the type the browser reports — both of those are attacker-influenced the moment someone is talked into converting a file they were sent. Anything that fails that check is rejected by name, with the reason, rather than producing a broken document.
Everything is assembled in memory in your browser, so the practical ceiling is your device rather than a server quota. A phone will reach it well before a laptop, and OCR in particular is real work: expect a second or two per clean page and longer for large or difficult ones.