ToolAct

EPUB to TXT Converter

Upload an EPUB e-book and extract its full text as a plain UTF-8 .txt file

Upload E-book

Drop an EPUB file here, or click to select

Supports .epub format (DRM-free), max 3MB

What is EPUB to TXT?

EPUB to TXT extracts the readable text out of an EPUB e-book and writes it as a plain text file. EPUB is a packaged format: the actual words live inside XHTML documents in a ZIP container, wrapped in CSS, fonts, images and navigation files. A .txt file keeps only the letters — no markup, no styling, no layout. What you get is a clean, continuous stream of text that any program can read.

This matters because plain text is the one format that everything understands: grep and full-text search, Python and R scripts for text mining, word counters, translation tools, LLM prompts, speech synthesis, and any device or application that has no e-reader support at all. If you want to search a library of books, quote passages precisely, or feed a novel into a language model, the first step is almost always the same — get the text out of the e-book container.

Understand the trade-off before you convert: TXT is a lossy destination. Images, the cover, the table of contents structure, bold and italic emphasis, tables, footnotes as linked notes, and embedded fonts all disappear or collapse into running text; only the characters survive, separated by line breaks and blank lines. Page numbers have no meaning either, since plain text has no pages. If you need the book still looking like a book, use an EPUB to PDF converter instead — TXT is the right choice when you need the words, not the pages.

How to Use

How to use

  1. Click the upload area or drag and drop an EPUB file — only .epub is supported, up to 3MB
  2. Click "Convert to TXT" and the server extracts the text with Calibre in a few seconds
  3. When the result panel shows the size change, click "Download TXT" to save the plain-text file
  4. Converting another book? Click "Convert Another File" to reset the panel

Before You Convert

  • Check the size change in the result panel: a large drop is normal because images and fonts are gone, but a result of only a few hundred bytes usually means the EPUB had almost no text layer.
  • Open the downloaded TXT in a UTF-8 aware editor once. If accented or CJK characters look wrong, your editor is falling back to a legacy encoding — the file itself is UTF-8.
  • If you need the book's layout, cover, or images, convert to PDF instead; TXT deliberately throws all of that away.

Use Cases

Full-text search across a bookReading a 600-page EPUB to find every mention of a character, term, or date is hopeless in an e-reader. Convert it to TXT and search it with grep, your editor, or a search engine — results come back instantly with line context, and you can open the same file on any machine without e-reader software.
Feed the text into a language model or NLP pipelineLLMs, summarizers, topic models, and corpus tools all take plain text. Removing the XHTML tags and CSS first avoids feeding the model markup noise that wastes context window and skews token counts. The UTF-8 output loads directly in Python with a single open() call, no EPUB library required.
Quote and cite accuratelyCopying from an e-reader app often drags along invisible formatting, smart-quote substitutions, and hyphenated line breaks. A TXT extraction gives you the raw characters as stored in the book, so quotes paste cleanly into a paper, a blog post, or a subtitle file — and you can verify the exact wording with a word count.
Text-to-speech and audio productionTTS engines and audiobook tools consume text, not EPUB. Extract the TXT first, split it into chapters by the blank-line or heading markers, then send it to the synthesiser. Stripping markup also stops readers from announcing image alt text or CSS leftovers.
Read on a device with no e-reader supportOld phones, e-ink readers with a locked-down firmware, a terminal, an embedded device, or a car head unit — plenty of screens can display a text file but cannot open an EPUB. A TXT copy is the lowest common denominator that always works, and it is small enough to email or sync anywhere.
Prepare a corpus for translation or subtitlingCAT tools and subtitle editors work on text segments. Doing a mechanical extraction first gives translators a stable, version-controllable source file; the segmentation step then happens once, on plain text, instead of fighting XHTML paragraph nesting inside the EPUB.

Technical Principle

The conversion is performed by Calibre's command-line tool ebook-convert, invoked on the server as ebook-convert <input.epub> <output.txt> with --txt-output-encoding utf-8 so the result is always UTF-8 regardless of the original declaration. The pipeline has three stages.

First, the EPUB is parsed. An EPUB is a ZIP archive whose first entry is an uncompressed mimetype file containing application/epub+zip, followed by META-INF/container.xml pointing at the root .opf package file. The .opf declares metadata (Dublin Core elements), a manifest listing every resource, and a spine defining reading order. The parser resolves the spine, then loads each XHTML content document in turn. This is why the text comes out in reading order rather than archive order — the spine is what defines the sequence, not the filenames.

Second, the XHTML is reduced to text. The DOM of each content document is walked and only text nodes are kept, so every element — <p>, <h1>, <em>, <span>, <div> — contributes its character data and nothing else. Block-level elements are converted into paragraph breaks, which is what produces the blank lines between paragraphs in the output; inline styling is simply dropped. Structural elements with no text content, such as <img> and <hr>, vanish entirely. Links lose their target but their anchor text remains in place, so URLs written as visible text survive while a linked word loses the address behind it.

Third, the result is serialised. Whitespace is normalised, the text is written out in the target encoding, and the file is finished. Two consequences are worth flagging. Any content that exists only as a picture is gone: a scanned novel, a comic in fixed-layout EPUB (rendition:layout=pre-paginated), or a children's picture book has essentially no text nodes, so the output is empty or nearly so. And no markup means no structure — CSS page-break rules, which the PDF pipeline uses to paginate, are meaningless here, so chapter boundaries survive only as whatever heading text the heading itself contained.

  • Engine: Calibre ebook-convert, run server-side as ebook-convert <in.epub> <out.txt> --txt-output-encoding utf-8
  • EPUB container: ZIP with an uncompressed mimetype (application/epub+zip), META-INF/container.xml pointing to the root .opf, and the .opf carrying metadata / manifest / spine
  • Reading order comes from the spine, not from filenames or ZIP order — this is why extracted chapters appear in the intended sequence
  • Text extraction keeps text nodes only: block elements become paragraph breaks, inline styling is discarded, and elements without character data (<img>, <hr>) disappear
  • Lossy by design: images, cover, tables (cells serialised as running text), fonts, colours, and TOC hierarchy do not survive; links keep only their anchor text
  • Image-only sources fail: a scanned EPUB or a fixed-layout comic has no text nodes, so the output is empty — the service detects the empty result and reports a failure instead of handing you a blank file
  • Output is UTF-8 plain text with normalised whitespace; no page concept, so page numbers and pagination settings have no effect

Examples

Search a whole book

Convert an EPUB to TXT, then grep -n "keyword" book.txt to jump straight to every mention with line numbers

Build an LLM prompt source

Extract a novel to TXT, split on blank lines, and load chapter by chapter into a summarisation or translation prompt

Count words and estimate reading time

wc -w book.txt after extraction gives a clean word count — impossible to get reliably from the EPUB's XHTML with markup

Prepare text for a TTS engine

Convert to TXT to strip the markup, then send the clean text to a speech synthesiser for audio production

FAQ

Where does my EPUB file go?

It is uploaded to the conversion server, processed with Calibre's ebook-convert, and returned to you as a .txt download. The file is not retained after the job finishes, but treat anything you upload as if it has left your device and avoid sending material you are not allowed to share. Your download link is a temporary task ID, not a permanent URL.

Will the images, cover, and formatting be kept?

No. TXT has no way to represent images, colours, fonts, or emphasis, so the converter keeps only the characters: bold and italic become plain text, the cover and inline figures are dropped, and CSS styling is discarded. Paragraph breaks are preserved as blank lines, which is the only structure plain text can carry.

Why is my TXT file empty, or why did the conversion fail?

Almost always because the book has no extractable text. Scanned EPUBs, comics, and fixed-layout picture books store each page as a picture, so there are no text nodes to extract. The service checks the output size and reports a conversion failure rather than giving you a zero-byte file. If your EPUB opens and displays fine in an e-reader but extracts to nothing, run OCR on the pages first.

Why is the TXT so much smaller than the EPUB?

Because text is the smallest part of an e-book. A typical EPUB carries the cover, inline illustrations, and embedded font files, all of which are compressed binary data; a novel's actual words are only a few hundred kilobytes. A drop of 80-95% is completely normal, and the result panel shows the exact ratio.

What about the table of contents and chapter headings?

The TOC is not regenerated as a TOC — TXT has no such concept — but the heading text itself usually survives in place, in reading order. nav.xhtml or toc.ncx entries are only kept if they also appear as visible text in the content. Practically, you can split the file on blank lines or on chapter-heading patterns to reconstruct sections.

Can I convert a DRM-protected EPUB?

No. Books locked with Adobe ADEPT, Apple FairPlay, or Amazon DRM cannot be opened by the converter — the content is encrypted and the parser fails before it reaches any text. Use a DRM-free EPUB you obtained legitimately, such as a self-made title, a public-domain work, or a file you have explicit permission to convert.

Which other e-book formats can be converted to TXT?

The server-side TXT endpoint accepts 15 input formats in total — epub, mobi, azw3, azw, fb2, lit, prc, pdb, html, htm, xhtml, rtf, odt, docx, and snb — but each format gets its own dedicated page on this site so the upload validation and the error messages match what you are actually converting. This page handles .epub only.

Is there a limit on how often I can convert?

Yes, the conversion endpoint is rate-limited to 2 requests per minute per IP address, counted across all e-book conversions. That is generous for normal use — it exists so a script cannot hammer the server — and waiting a minute clears it if you hit the limit.

What encoding is the output, and will CJK text work?

Output is always UTF-8, forced by the --txt-output-encoding utf-8 argument, so Chinese, Japanese, Korean, Cyrillic, Greek, and accented Latin text all come through intact. If the characters render as garbage after download, the fault is the editor's detected encoding, not the file — switch it to UTF-8 manually.