Skip to content

DOI Extractor

Extract, clean and deduplicate DOIs from reference lists and flag malformed ones.

DOIs are found in plain text, doi: prefixes and doi.org links (including URL-encoded ones). Trailing punctuation is removed while balanced brackets, as in 10.1016/S0140-6736(20)30183-5, are kept. Everything runs in your browser.

Continue your work

Share this tool

Found it useful? Send it to a friend or teammate.

Rate this tool

0.0

0 ratings

  • 5 stars 0
  • 4 stars 0
  • 3 stars 0
  • 2 stars 0
  • 1 star 0

Click a star to rate this tool

Clear instructions

Find steps, examples and limitations below.

Use online

Open the tool in a supported web browser.

Free to use

No sign-up required. Tool-specific limits may apply.

How to use the DOI Extractor

Extract, clean and deduplicate DOIs from reference lists and flag malformed ones.

  1. 1 Paste your reference list, bibliography or any text.
  2. 2 Choose whether to remove duplicates and how to format the output.
  3. 3 Copy the DOIs or download them as TXT or CSV, and review any flagged lines.

Example and practical tips

“…Lancet. 2020;395:497-506. https://doi.org/10.1016/S0140-6736(20)30183-5.” gives 10.1016/S0140-6736(20)30183-5, keeping the brackets and dropping the final full stop.

Frequently asked questions

Are DOIs case-sensitive?

No. 10.1136/BMJ.N71 and 10.1136/bmj.n71 are the same DOI, so duplicates are matched without regard to case.

Why is a line flagged?

It mentions a DOI that could not be read, often because it was split across two lines or is missing the 10. prefix.

Report an issue

Something broken or not quite right? Tell us and we will look into it.