Extract Text

Extract all text from a PDF as a plain text file

pdfty's extract tool pulls every character of text out of a PDF and delivers it as a clean plain-text file. It beats select-all copy-paste, which routinely scrambles multi-column layouts, drops text boxes, and chokes on long documents โ€” extraction processes all 50 pages in one pass with reading order preserved.

Extraction runs on PyMuPDF, which reads the PDF's internal text structure directly and reconstructs natural reading order โ€” including two-column layouts, tables, headers, and footnotes. Full Unicode comes through intact: Cyrillic, CJK, Arabic, accented characters. The result is ideal for feeding documents into translation tools, search indexes, or AI assistants, for quoting from papers, and for word counts.

One important note: this works on PDFs that contain a text layer. Scanned documents are just pictures of text โ€” run them through our OCR tool (powered by Tesseract) first, then extract. Free for files up to 20 MB and 50 pages, no watermark, no signup. Files are permanently deleted within 1 hour.

How it works

  1. 1

    Upload your PDF

    Drag and drop or click to select. Up to 20 MB and 50 pages free.

  2. 2

    Extract

    PyMuPDF reads the text layer and reconstructs reading order. A second or two for most files.

  3. 3

    Preview

    Check the first pages of extracted text on screen.

  4. 4

    Download the .txt

    One plain-text file, with page breaks marked.

  5. 5

    Done in 1 hour

    Your PDF and the text file are deleted from our servers within 1 hour.

Frequently asked questions

Your PDF is almost certainly a scan โ€” images of text with no text layer. Run it through our OCR tool first, then extract.

Related tools