EPUB to TXT Converter
Upload an EPUB e-book and extract its full text as a plain UTF-8 .txt file
Drop an EPUB file here, or click to select
Supports .epub format (DRM-free), max 3MB
What is EPUB to TXT?
EPUB to TXT extracts the readable text out of an EPUB e-book and writes it as a plain text file. EPUB is a packaged format: the actual words live inside XHTML documents in a ZIP container, wrapped in CSS, fonts, images and navigation files. A .txt file keeps only the letters — no markup, no styling, no layout. What you get is a clean, continuous stream of text that any program can read.
This matters because plain text is the one format that everything understands: grep and full-text search, Python and R scripts for text mining, word counters, translation tools, LLM prompts, speech synthesis, and any device or application that has no e-reader support at all. If you want to search a library of books, quote passages precisely, or feed a novel into a language model, the first step is almost always the same — get the text out of the e-book container.
Understand the trade-off before you convert: TXT is a lossy destination. Images, the cover, the table of contents structure, bold and italic emphasis, tables, footnotes as linked notes, and embedded fonts all disappear or collapse into running text; only the characters survive, separated by line breaks and blank lines. Page numbers have no meaning either, since plain text has no pages. If you need the book still looking like a book, use an EPUB to PDF converter instead — TXT is the right choice when you need the words, not the pages.
How to Use
How to use
- Click the upload area or drag and drop an EPUB file — only .epub is supported, up to 3MB
- Click "Convert to TXT" and the server extracts the text with Calibre in a few seconds
- When the result panel shows the size change, click "Download TXT" to save the plain-text file
- Converting another book? Click "Convert Another File" to reset the panel
Before You Convert
- Check the size change in the result panel: a large drop is normal because images and fonts are gone, but a result of only a few hundred bytes usually means the EPUB had almost no text layer.
- Open the downloaded TXT in a UTF-8 aware editor once. If accented or CJK characters look wrong, your editor is falling back to a legacy encoding — the file itself is UTF-8.
- If you need the book's layout, cover, or images, convert to PDF instead; TXT deliberately throws all of that away.
Use Cases
Technical Principle
The conversion is performed by Calibre's command-line tool ebook-convert, invoked on the server as ebook-convert <input.epub> <output.txt> with --txt-output-encoding utf-8 so the result is always UTF-8 regardless of the original declaration. The pipeline has three stages.
First, the EPUB is parsed. An EPUB is a ZIP archive whose first entry is an uncompressed mimetype file containing application/epub+zip, followed by META-INF/container.xml pointing at the root .opf package file. The .opf declares metadata (Dublin Core elements), a manifest listing every resource, and a spine defining reading order. The parser resolves the spine, then loads each XHTML content document in turn. This is why the text comes out in reading order rather than archive order — the spine is what defines the sequence, not the filenames.
Second, the XHTML is reduced to text. The DOM of each content document is walked and only text nodes are kept, so every element — <p>, <h1>, <em>, <span>, <div> — contributes its character data and nothing else. Block-level elements are converted into paragraph breaks, which is what produces the blank lines between paragraphs in the output; inline styling is simply dropped. Structural elements with no text content, such as <img> and <hr>, vanish entirely. Links lose their target but their anchor text remains in place, so URLs written as visible text survive while a linked word loses the address behind it.
Third, the result is serialised. Whitespace is normalised, the text is written out in the target encoding, and the file is finished. Two consequences are worth flagging. Any content that exists only as a picture is gone: a scanned novel, a comic in fixed-layout EPUB (rendition:layout=pre-paginated), or a children's picture book has essentially no text nodes, so the output is empty or nearly so. And no markup means no structure — CSS page-break rules, which the PDF pipeline uses to paginate, are meaningless here, so chapter boundaries survive only as whatever heading text the heading itself contained.
- Engine: Calibre ebook-convert, run server-side as ebook-convert <in.epub> <out.txt> --txt-output-encoding utf-8
- EPUB container: ZIP with an uncompressed mimetype (application/epub+zip), META-INF/container.xml pointing to the root .opf, and the .opf carrying metadata / manifest / spine
- Reading order comes from the spine, not from filenames or ZIP order — this is why extracted chapters appear in the intended sequence
- Text extraction keeps text nodes only: block elements become paragraph breaks, inline styling is discarded, and elements without character data (<img>, <hr>) disappear
- Lossy by design: images, cover, tables (cells serialised as running text), fonts, colours, and TOC hierarchy do not survive; links keep only their anchor text
- Image-only sources fail: a scanned EPUB or a fixed-layout comic has no text nodes, so the output is empty — the service detects the empty result and reports a failure instead of handing you a blank file
- Output is UTF-8 plain text with normalised whitespace; no page concept, so page numbers and pagination settings have no effect
Examples
Search a whole book
Convert an EPUB to TXT, then grep -n "keyword" book.txt to jump straight to every mention with line numbersBuild an LLM prompt source
Extract a novel to TXT, split on blank lines, and load chapter by chapter into a summarisation or translation promptCount words and estimate reading time
wc -w book.txt after extraction gives a clean word count — impossible to get reliably from the EPUB's XHTML with markupPrepare text for a TTS engine
Convert to TXT to strip the markup, then send the clean text to a speech synthesiser for audio productionFAQ
Where does my EPUB file go?
It is uploaded to the conversion server, processed with Calibre's ebook-convert, and returned to you as a .txt download. The file is not retained after the job finishes, but treat anything you upload as if it has left your device and avoid sending material you are not allowed to share. Your download link is a temporary task ID, not a permanent URL.
Will the images, cover, and formatting be kept?
No. TXT has no way to represent images, colours, fonts, or emphasis, so the converter keeps only the characters: bold and italic become plain text, the cover and inline figures are dropped, and CSS styling is discarded. Paragraph breaks are preserved as blank lines, which is the only structure plain text can carry.
Why is my TXT file empty, or why did the conversion fail?
Almost always because the book has no extractable text. Scanned EPUBs, comics, and fixed-layout picture books store each page as a picture, so there are no text nodes to extract. The service checks the output size and reports a conversion failure rather than giving you a zero-byte file. If your EPUB opens and displays fine in an e-reader but extracts to nothing, run OCR on the pages first.
Why is the TXT so much smaller than the EPUB?
Because text is the smallest part of an e-book. A typical EPUB carries the cover, inline illustrations, and embedded font files, all of which are compressed binary data; a novel's actual words are only a few hundred kilobytes. A drop of 80-95% is completely normal, and the result panel shows the exact ratio.
What about the table of contents and chapter headings?
The TOC is not regenerated as a TOC — TXT has no such concept — but the heading text itself usually survives in place, in reading order. nav.xhtml or toc.ncx entries are only kept if they also appear as visible text in the content. Practically, you can split the file on blank lines or on chapter-heading patterns to reconstruct sections.
Can I convert a DRM-protected EPUB?
No. Books locked with Adobe ADEPT, Apple FairPlay, or Amazon DRM cannot be opened by the converter — the content is encrypted and the parser fails before it reaches any text. Use a DRM-free EPUB you obtained legitimately, such as a self-made title, a public-domain work, or a file you have explicit permission to convert.
Which other e-book formats can be converted to TXT?
The server-side TXT endpoint accepts 15 input formats in total — epub, mobi, azw3, azw, fb2, lit, prc, pdb, html, htm, xhtml, rtf, odt, docx, and snb — but each format gets its own dedicated page on this site so the upload validation and the error messages match what you are actually converting. This page handles .epub only.
Is there a limit on how often I can convert?
Yes, the conversion endpoint is rate-limited to 2 requests per minute per IP address, counted across all e-book conversions. That is generous for normal use — it exists so a script cannot hammer the server — and waiting a minute clears it if you hit the limit.
What encoding is the output, and will CJK text work?
Output is always UTF-8, forced by the --txt-output-encoding utf-8 argument, so Chinese, Japanese, Korean, Cyrillic, Greek, and accented Latin text all come through intact. If the characters render as garbage after download, the fault is the editor's detected encoding, not the file — switch it to UTF-8 manually.
Related Tools
EPUB to PDF Converter
Free online EPUB to PDF converter. Turn EPUB e-books into layout-stable PDFs with chapters, images and fonts preserved. Files auto-delete after processing.
MOBI to EPUB Converter
Free online MOBI to EPUB converter. Turn Kindle MOBI e-books into the universal reflowable EPUB format with chapters, cover and images preserved. Files auto-delete after processing.
MOBI to PDF Converter
Free online MOBI to PDF converter. Turn Kindle MOBI e-books into layout-stable PDFs with chapters, images and text preserved. Files auto-delete after processing.
AZW3 to PDF Converter
Free online AZW3 to PDF converter. Turn Kindle AZW3 (KF8) e-books into print-ready PDFs with chapters, images and fonts preserved. Files auto-delete after processing.
PDF to Word Converter
Free online PDF to Word converter. Export .docx or .doc files with layout and image preservation, one-click conversion, and automatic file deletion.
Word to PDF Converter
Free online Word to PDF converter for .docx and .doc files. Convert documents with layout preservation, optimized PDF output, and automatic file deletion.
