Open the citation extractor

Recover citations from dead bibliographies

Recover a Citation Library From Word, DOCX, or Plain Text References

If the EndNote, Zotero, Mendeley, RefWorks, Paperpile, or library-manager database is gone, but the bibliography still exists, LumaCite can rebuild structured citation data from Word, DOCX, Google Docs, TXT, RTF, RIS, BibTeX, CSV, or pasted references.

Lost EndNote library Word bibliography to RIS DOCX references to BibTeX Plain text references to Zotero DOI and PMID recovery

Reference manager recovery

Rebuild a citation library when all you have left is the bibliography.

Many researchers still have the manuscript, thesis, systematic review, dissertation, grant draft, or final DOCX file, but the live citation database is unavailable. This is where bibliography extraction is a real advantage: the document may look static, but the references still contain titles, authors, journals, years, DOI strings, PMID links, and enough metadata evidence to reconstruct useful records.

Lost EndNote library Recover references from the final Word bibliography when the `.enl` file, data folder, or traveling library is unavailable.
Plain text manuscript Extract citations from TXT, Markdown, email, LaTeX notes, pasted manuscript references, or archived web copy.
DOCX without live citations Upload a Word document whose bibliography is no longer connected to a citation manager.
Team handoff cleanup Turn a collaborator's pasted reference list into RIS, BibTeX, CSL-JSON, CSV, or an audit report.
Systematic review repair Check hundreds of pasted rows for missing years, duplicate references, DOI/PMID coverage, and metadata source agreement.
Migration to a new manager Move a dead bibliography into Zotero, EndNote, Mendeley, Overleaf, Word, Google Docs, or a local library manager.

How LumaCite compares with common citation recovery options

Workflow Good for Typical gap LumaCite advantage
Desktop reference manager import Already structured RIS, BibTeX, CSL, or live library files Usually cannot rebuild clean records from a dead formatted bibliography alone Starts from Word, DOCX, Google Docs copy, and plain text reference lists
DOI citation generator One known DOI, PMID, ISBN, arXiv ID, or URL at a time Slow for 50, 100, or 300 references and weak when many rows have no DOI Batch-splits the bibliography, extracts identifiers, and enriches many rows
PDF reference extractor References locked inside a paper PDF PDF layout can scramble columns, line order, and reference boundaries Uses cleaner Word/text structure when the manuscript or bibliography text is available
Manual search and copy Checking a few uncertain references Manual, inconsistent, and easy to import wrong editions or duplicate records Combines parsing, DOI/PMID detection, metadata checks, duplicates, and exports
Regex or spreadsheet cleanup Very regular numbered lists Fails on Vancouver rows, author-year rows, merged rows, missing punctuation, and mixed formats Uses multiple boundary signals, count reconciliation, source evidence, and review states

The engine this page promotes

The actual extractor lives on the focused tool page.

This page explains the recovery use case. The working engine is the Extract Citations From Text tool, where users upload DOCX, paste references, run metadata checks, review warnings, and download citation exports.

From dead text to live citation data

Use one recovered bibliography across your writing stack.

Extract references from text once, then move the cleaned records into the tool that fits the next stage: Zotero for research organization, EndNote for Word manuscripts, BibTeX for LaTeX and Overleaf, CSL-JSON for modern citation processors, CSV for audits, and Markdown for clean review notes.

Word bibliography to RIS Word bibliography to BibTeX DOCX references to CSL-JSON Plain text references to Zotero Google Docs references to CSV Vancouver references to BibTeX APA reference list parser DOI extractor from bibliography PMID extractor from references Reference cleanup audit Lost EndNote library recovery Manuscript bibliography repair
Count checks The engine compares extracted rows with independent signals such as PubMed links, DOI counts, numbered markers, and paragraph starts.
Boundary repair Rows with repeated years, PubMed markers, DOI strings, or author restarts are flagged and repaired before export.
Biomedical fallback Rows without DOI or PMID can still be searched against PubMed and Crossref, then scored by title, year, source, and agreement.
Audit-first exports Download an audit report with raw text, confidence, warnings, identifiers, duplicate risk, and source checks.
lost EndNote library recovery recover references from Word bibliography Word bibliography to Zotero Word bibliography to EndNote extract citations from text bibliography parser reference list parser text to BibTeX text to RIS text to CSL-JSON extract references from Word extract citations from Google Docs DOCX bibliography extractor DOI extractor from references PMID extractor from bibliography

Search questions this page answers

Citation extraction Q&A for Word, DOCX, TXT, RIS, BibTeX, and lost libraries

How do I recover references if I lost my EndNote library?

Upload the DOCX or paste the bibliography in the text citation extractor. LumaCite extracts each citation row, detects DOI/PMID/URL evidence, enriches metadata, and exports RIS or BibTeX so you can rebuild the library.

Can I convert plain text references to Zotero?

Yes. Paste the plain text reference list, review the parsed rows, then export RIS, BibTeX, CSL-JSON, or CSV for import and cleanup.

Can I extract references from a DOCX file?

Yes. The extractor reads DOCX text and hyperlink targets, including PubMed links when present, before running the bibliography parser.

Can I rebuild citations from a PDF instead?

Use the dedicated PDF reference extractor for paper PDFs. Use the text extractor when you have cleaner text from Word, Google Docs, TXT, RTF, RIS, BibTeX, or CSV.

Can this handle mixed citation styles?

It is designed for numbered references, bracketed references, Vancouver-style biomedical rows, author-year rows, blank-line sections, RIS tags, and BibTeX entries.

What makes this better than manual citation lookup?

The tool combines batch extraction, identifier detection, metadata lookup, duplicate detection, count checks, editable review, and export files in one workflow.

Use the engine

Ready to recover the bibliography?

Where do I actually upload the Word document or paste references?

Use the focused Extract Citations From Text page. This page exists to explain and promote the recovery workflow.

What should I submit in Google Search Console?

Submit both this SEO page and the extractor page, then resubmit the sitemap after deployment.

What exports does the extractor create?

BibTeX, RIS, CSL-JSON, CSV, Markdown, and an audit report with confidence, warnings, raw text, and source checks.

Open the extractor