Extract citations from a PDF or text

Pull every citation out of a PDF or pasted text — full bibliographic entries from the references section, in-text markers (Smith 2023; [14]) mapped to their source, and any DOIs or URLs. Export as BibTeX, RIS, CSL JSON, or a flat spreadsheet.

Drop a PDF or image here, or browse
PDF or image · up to 20 MB
Processed in-flight — never stored on our servers.

What should we pull from this document?

Or pick specific fields

Or describe it yourself

Why this matters

Bibliography management tools (Zotero, Mendeley, EndNote) extract references from PDFs by pattern-matching on font and layout — they break on legal briefs, government reports, and any document that doesn't follow APA/MLA conventions. ExtractFox reads the page semantically, so it picks up citations regardless of style and reliably maps in-text markers to their full reference. Systematic reviews require hundreds of references formatted consistently — manually copying author names from a thesis bibliography with mixed comma and ampersand styles takes hours and invites typos in DOIs. Litigation teams citing 200+ cases in a brief need every Bluebook citation checked against the references section, but numbered footnotes and pin cites don't match what Zotero's regex expects.

How it works

  1. Step 1
    Upload the PDF or paste text

    Research papers, legal briefs, reports, or pasted text. Multi-column layouts and footnoted citations both work.

  2. Step 2
    Pick a mode

    All citations, in-text markers only, references section only, or grouped by section — and choose an export format.

  3. Step 3
    Export

    BibTeX (.bib), RIS, CSL JSON, or CSV. Drop straight into Zotero, Mendeley, EndNote, or a spreadsheet.

Common use cases

Systematic review prep — pull every reference from 50+ PDFs into one Zotero library
Litigation briefs — extract Bluebook case citations with reporter, volume, and pinpoint pages
Literature review spreadsheets — author, year, title, DOI as flat rows for screening
Thesis bibliography audit — find missing DOIs and incomplete entries before submission

Sample output

Example: 3 of the references from a research paper PDF, exported as CSL JSON

idtypetitleauthorissuedcontainer-titlevolumeDOI
vaswani2017attentionpaper-conferenceAttention Is All You Need[{"family":"Vaswani","given":"Ashish"},{"family":"Shazeer","given":"Noam"},{"family":"Parmar","given":"Niki"}]{"date-parts":[[2017]]}Advances in Neural Information Processing Systems30
devlin2019bertpaper-conferenceBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding[{"family":"Devlin","given":"Jacob"}]{"date-parts":[[2019]]}NAACL-HLT10.18653/v1/N19-1423

Frequently asked questions

How do I extract citations from a PDF?+

Upload the PDF, pick a mode (all citations, in-text mapped, BibTeX, CSL JSON, or legal citations), and export. The result drops straight into Zotero, Mendeley, EndNote, or any reference manager that accepts BibTeX or RIS.

How is this different from Zotero's PDF metadata extraction?+

Zotero's extractor reads metadata stored in the PDF and pattern-matches the references section. It works well on standard APA/MLA papers and breaks on most legal briefs, government reports, and theses. ExtractFox reads the page semantically — it works regardless of citation style.

Will it map in-text markers to their references?+

Yes — pick the In-text markers mapped mode. Each marker comes back with the sentence it appeared in, the page, and the full reference it points to.

What about footnoted citations?+

Handled the same way. Citations in footnotes are tagged with the page they appeared on; the full bibliographic version still ends up in the bibliography list.

What citation styles does it recognize?+

APA, MLA, Chicago, IEEE, Vancouver, Bluebook, ACS, AMA, and most journal-specific variants. The mode detects the style and parses fields accordingly.

Can it extract citations from a thesis or dissertation with mixed bibliography styles?+

Yes. Theses often mix APA body citations with a Chicago bibliography or append references from multiple projects. ExtractFox reads each entry semantically rather than matching one style template.

Does it handle numbered citations like [1] and superscript footnotes?+

Yes — pick In-text markers mapped mode. Numbered brackets, superscripts, and author-year parentheticals all map back to the full bibliographic entry in the references section.

Related extractors