Getting usable text out of PDFs and scans
Text tools work best when you first know what kind of PDF you have. Some files contain real selectable text, while scanned documents may only contain page images. OCR is useful for the second case, but recognition quality depends heavily on scan sharpness, page orientation, language, and the type of material on the page. Tables, multi-column layouts, handwriting, and faint copies usually need more checking than a clean single-column document.
After extraction, do not judge the result only by whether words appeared. Look for missing lines, repeated headers, broken hyphenation, wrong characters, and sections that were read in the wrong order. If the text will be quoted, searched, or reused in another document, compare important passages against the original page before treating the extracted version as authoritative.