Recover citations from dead bibliographies
Recover a Citation Library From Word, DOCX, or Plain Text References
If the EndNote, Zotero, Mendeley, RefWorks, Paperpile, or library-manager database is gone, but the bibliography still exists, LumaCite can rebuild structured citation data from Word, DOCX, Google Docs, TXT, RTF, RIS, BibTeX, CSV, or pasted references.
Reference manager recovery
Rebuild a citation library when all you have left is the bibliography.
Many researchers still have the manuscript, thesis, systematic review, dissertation, grant draft, or final DOCX file, but the live citation database is unavailable. This is where bibliography extraction is a real advantage: the document may look static, but the references still contain titles, authors, journals, years, DOI strings, PMID links, and enough metadata evidence to reconstruct useful records.
How LumaCite compares with common citation recovery options
| Workflow | Good for | Typical gap | LumaCite advantage |
|---|---|---|---|
| Desktop reference manager import | Already structured RIS, BibTeX, CSL, or live library files | Usually cannot rebuild clean records from a dead formatted bibliography alone | Starts from Word, DOCX, Google Docs copy, and plain text reference lists |
| DOI citation generator | One known DOI, PMID, ISBN, arXiv ID, or URL at a time | Slow for 50, 100, or 300 references and weak when many rows have no DOI | Batch-splits the bibliography, extracts identifiers, and enriches many rows |
| PDF reference extractor | References locked inside a paper PDF | PDF layout can scramble columns, line order, and reference boundaries | Uses cleaner Word/text structure when the manuscript or bibliography text is available |
| Manual search and copy | Checking a few uncertain references | Manual, inconsistent, and easy to import wrong editions or duplicate records | Combines parsing, DOI/PMID detection, metadata checks, duplicates, and exports |
| Regex or spreadsheet cleanup | Very regular numbered lists | Fails on Vancouver rows, author-year rows, merged rows, missing punctuation, and mixed formats | Uses multiple boundary signals, count reconciliation, source evidence, and review states |
The engine this page promotes
The actual extractor lives on the focused tool page.
This page explains the recovery use case. The working engine is the Extract Citations From Text tool, where users upload DOCX, paste references, run metadata checks, review warnings, and download citation exports.
From dead text to live citation data
Use one recovered bibliography across your writing stack.
Extract references from text once, then move the cleaned records into the tool that fits the next stage: Zotero for research organization, EndNote for Word manuscripts, BibTeX for LaTeX and Overleaf, CSL-JSON for modern citation processors, CSV for audits, and Markdown for clean review notes.
Search questions this page answers
Citation extraction Q&A for Word, DOCX, TXT, RIS, BibTeX, and lost libraries
How do I recover references if I lost my EndNote library?
Upload the DOCX or paste the bibliography in the text citation extractor. LumaCite extracts each citation row, detects DOI/PMID/URL evidence, enriches metadata, and exports RIS or BibTeX so you can rebuild the library.
Can I convert plain text references to Zotero?
Yes. Paste the plain text reference list, review the parsed rows, then export RIS, BibTeX, CSL-JSON, or CSV for import and cleanup.
Can I extract references from a DOCX file?
Yes. The extractor reads DOCX text and hyperlink targets, including PubMed links when present, before running the bibliography parser.
Can I rebuild citations from a PDF instead?
Use the dedicated PDF reference extractor for paper PDFs. Use the text extractor when you have cleaner text from Word, Google Docs, TXT, RTF, RIS, BibTeX, or CSV.
Can this handle mixed citation styles?
It is designed for numbered references, bracketed references, Vancouver-style biomedical rows, author-year rows, blank-line sections, RIS tags, and BibTeX entries.
What makes this better than manual citation lookup?
The tool combines batch extraction, identifier detection, metadata lookup, duplicate detection, count checks, editable review, and export files in one workflow.
Use the engine
Ready to recover the bibliography?
Where do I actually upload the Word document or paste references?
Use the focused Extract Citations From Text page. This page exists to explain and promote the recovery workflow.
What should I submit in Google Search Console?
Submit both this SEO page and the extractor page, then resubmit the sitemap after deployment.
What exports does the extractor create?
BibTeX, RIS, CSL-JSON, CSV, Markdown, and an audit report with confidence, warnings, raw text, and source checks.