MarkdownPDF

Research papers to Markdown, columns and tables included

A paper is laid out for print: two columns, floating tables, figures, footnotes. Copying its text gives you sentences in the wrong order. Converted properly, it becomes notes you can search, quote and feed to the tools you use.

The two-column problem

A PDF stores text by position, not by reading order. In a two-column paper, a naive extraction reads straight across the page and interleaves the columns line by line. The converter finds the gutter between columns and reads each column top to bottom; titles, wide figures and tables that span both columns are kept in place.

Tables are rebuilt as Markdown tables when their cells line up — results tables usually do. Figures can be extracted as image files and linked where they appeared. Equations stay as the characters the PDF contains: readable, but not LaTeX.

Where the Markdown goes next

Into Obsidian or Logseq for a literature review, with properties recording the source and page markers for citations. Into ChatGPT, Claude or NotebookLM, where clean structure means better answers and fewer tokens than raw PDF text. Into a chunker, for a retrieval pipeline over a corpus of papers.

A typical workflow

  1. 1

    Convert the paper

    Check the result against the original side by side — click any line to find it in the PDF.

    Open PDF to Markdown →
  2. 2

    File it in your vault

    A note with title, source and page count as properties, figures as attachments, and hidden page markers for citing.

    Open PDF to Obsidian →
  3. 3

    Do a whole reading list

    Drop every PDF of the literature review at once and get one ZIP of Markdown files.

    Open Batch PDF to Markdown →
  4. 4

    Pull out the results tables

    Export the tables of a paper to CSV to re-plot or compare them.

    Open PDF Tables to CSV →

Frequently asked questions

Still stuck? Ask us.

Does it handle equations?

Inline symbols and simple formulas come through as text (α, β, R + 1). Display equations built from positioned glyphs often lose their layout. If equations matter most, a dedicated math-OCR tool will do better; the comparison page lists the honest alternatives.

What about references and footnotes?

They are converted as text in reading order. Footnote markers stay as plain numbers next to the word they annotate — they are not turned into Markdown footnote links.

Can I convert a scanned old article?

Yes. Scanned pages go through OCR in your browser and are flagged for review in the result.

Tools for this job

Other use cases: Lawyers & legal teams · Finance & accounting · Students