Research papers to Markdown, columns and tables included
A paper is laid out for print: two columns, floating tables, figures, footnotes. Copying its text gives you sentences in the wrong order. Converted properly, it becomes notes you can search, quote and feed to the tools you use.
The two-column problem
A PDF stores text by position, not by reading order. In a two-column paper, a naive extraction reads straight across the page and interleaves the columns line by line. The converter finds the gutter between columns and reads each column top to bottom; titles, wide figures and tables that span both columns are kept in place.
Tables are rebuilt as Markdown tables when their cells line up — results tables usually do. Figures can be extracted as image files and linked where they appeared. Equations stay as the characters the PDF contains: readable, but not LaTeX.
Where the Markdown goes next
Into Obsidian or Logseq for a literature review, with properties recording the source and page markers for citations. Into ChatGPT, Claude or NotebookLM, where clean structure means better answers and fewer tokens than raw PDF text. Into a chunker, for a retrieval pipeline over a corpus of papers.
A typical workflow
- 1
Convert the paper
Check the result against the original side by side — click any line to find it in the PDF.
Open PDF to Markdown → - 2
File it in your vault
A note with title, source and page count as properties, figures as attachments, and hidden page markers for citing.
Open PDF to Obsidian → - 3
Do a whole reading list
Drop every PDF of the literature review at once and get one ZIP of Markdown files.
Open Batch PDF to Markdown → - 4
Pull out the results tables
Export the tables of a paper to CSV to re-plot or compare them.
Open PDF Tables to CSV →
Frequently asked questions
Still stuck? Ask us.
Does it handle equations?
Inline symbols and simple formulas come through as text (α, β, R + 1). Display equations built from positioned glyphs often lose their layout. If equations matter most, a dedicated math-OCR tool will do better; the comparison page lists the honest alternatives.
What about references and footnotes?
They are converted as text in reading order. Footnote markers stay as plain numbers next to the word they annotate — they are not turned into Markdown footnote links.
Can I convert a scanned old article?
Yes. Scanned pages go through OCR in your browser and are flagged for review in the result.
Tools for this job
- PDF to MarkdownClean Markdown from any PDF, with OCR for scanned pages.
- PDF to ObsidianA PDF as an Obsidian note: properties, page markers, embedded images.
- Batch PDF to MarkdownA whole folder of PDFs to Markdown in one go, as a ZIP.
- PDF Tables to CSVEvery table in a PDF, spreadsheet-ready — one CSV per table.
- Markdown Chunker for RAGSplit documents into token-sized chunks that respect headings.
Other use cases: Lawyers & legal teams · Finance & accounting · Students