DOI Extractor
Extract, clean and deduplicate DOIs from reference lists and flag malformed ones.
DOIs are found in plain text, doi: prefixes and doi.org links (including URL-encoded ones). Trailing punctuation is removed while balanced brackets, as in 10.1016/S0140-6736(20)30183-5, are kept. Everything runs in your browser.
Continue your work
- Reference Formatter (DOI to Citation) — Turn the DOIs into formatted references.
- PMID Extractor — Extract PubMed IDs as well.
Share this tool
Found it useful? Send it to a friend or teammate.
Rate this tool
0.0
0 ratings
- 5 stars 0
- 4 stars 0
- 3 stars 0
- 2 stars 0
- 1 star 0
Clear instructions
Find steps, examples and limitations below.
Use online
Open the tool in a supported web browser.
Free to use
No sign-up required. Tool-specific limits may apply.
How to use the DOI Extractor
Extract, clean and deduplicate DOIs from reference lists and flag malformed ones.
- 1 Paste your reference list, bibliography or any text.
- 2 Choose whether to remove duplicates and how to format the output.
- 3 Copy the DOIs or download them as TXT or CSV, and review any flagged lines.
Example and practical tips
“…Lancet. 2020;395:497-506. https://doi.org/10.1016/S0140-6736(20)30183-5.” gives 10.1016/S0140-6736(20)30183-5, keeping the brackets and dropping the final full stop.
Frequently asked questions
Are DOIs case-sensitive?
No. 10.1136/BMJ.N71 and 10.1136/bmj.n71 are the same DOI, so duplicates are matched without regard to case.
Why is a line flagged?
It mentions a DOI that could not be read, often because it was split across two lines or is missing the 10. prefix.
Report an issue
Something broken or not quite right? Tell us and we will look into it.