One API call converts resumes in PDF, DOCX, scanned images and five other formats into 114+ structured JSON fields. Three normalization APIs standardize job titles, skills and locations. Built-in anonymization removes candidate identity for bias-free screening. From $0.0198 per parse with multilingual OCR support.
114+ fields per resume
Convert any resume file into a structured JSON candidate profile. AI-powered parsing with OCR, seniority computation, skill normalization and built-in anonymization for bias-free screening workflows.
Explore platformResumes arrive in dozens of formats, use inconsistent terminology and contain sensitive personal information that must be handled carefully. Manual processing is slow, expensive and introduces errors and bias at every step of the recruitment workflow.
PDF documents with complex layouts. DOCX files with embedded tables and columns. Scanned paper resumes saved as image-based PDFs. Rich text files from legacy systems. HTML exports from LinkedIn and job boards. Even XLSX spreadsheets from staffing agencies. The Resume Reader API accepts PDF, DOCX, TXT, RTF, HTML, ODT, XLSX and XLS files up to 10 MB and 20 pages each, with built-in OCR for scanned and image-based documents.
Recruiters spend hours copying candidate information from documents into applicant tracking systems. Every manual keystroke introduces the possibility of transposition errors, missed fields and inconsistent formatting. At scale, a team processing hundreds of resumes per day loses significant time to data entry that could be spent on candidate evaluation and relationship building.
The same role appears as "Sr. SWE II", "Senior Software Engineer", "Staff Developer" and "Lead Programmer" across different resumes. Skills appear as "K8s", "Kubernetes", "k8s orchestration" and "container orchestration". Locations read as "NYC", "New York City", "New York, NY" and "Manhattan". Without normalization, matching, filtering and reporting break down. The three normalization APIs resolve 100% of inputs with deterministic accuracy rates of 95.4%, 98.2% and 94.5% respectively.
Research consistently shows that candidate names, ages, gender indicators and addresses influence hiring decisions. Organizations pursuing equitable hiring practices need a reliable way to remove identifying information from resumes before human review. The Resume Reader API includes a built-in anonymization parameter that strips all personal identifiers while preserving the skills, experience and qualifications that matter for evaluation.
The Resume Reader API converts unstructured document content into a comprehensive structured JSON profile. Every field is evidence-faithful: if the information is not present in the source document, the API returns null rather than fabricating data. These are the core field groups returned per candidate.
Raw resume data is inconsistent by nature. The same job title, skill or location appears in dozens of variations across candidates. The three normalization APIs resolve 100% of inputs into standardized forms at 0.1 credits per item, enabling consistent matching, filtering and reporting across your entire candidate database.
"Sr. SWE II" becomes "Senior Software Engineer". "VP of Eng" becomes "Vice President of Engineering". "Dev Lead" becomes "Development Lead". The job title normalization API maps every variation, abbreviation and informal title to a canonical form. 100% of inputs resolved with 95.4% deterministic accuracy.
Each normalized title is tagged with a seniority level: intern, junior, mid-level, senior, lead, principal, director, vice president, C-suite. The classification enables experience-based routing and compensation benchmarking across your entire candidate pool.
Job titles vary dramatically across industries. A "Producer" in media is different from a "Producer" in manufacturing. The normalization engine uses context from the resume's industry signals and company information to disambiguate titles and assign the correct canonical form and seniority level.
"K8s" becomes "Kubernetes". "Phyton" becomes "Python". "JS" becomes "JavaScript". "ML" becomes "Machine Learning". "TF" becomes "TensorFlow". The skill normalization API handles abbreviations, misspellings, acronyms and informal names. 100% resolution rate with 98.2% deterministic accuracy across the full technology and business skill landscape.
"React.js", "ReactJS", "React" and "React Framework" all map to a single canonical "React" entry. "PostgreSQL", "Postgres" and "PG" converge to "PostgreSQL". Synonym clustering eliminates duplicate skill entries and enables accurate skill frequency analysis across candidate pools.
Each normalized skill is tagged with its category: programming language, framework, database, cloud platform, methodology, soft skill and dozens more. Category tags enable structured skill inventory reports and gap analysis for workforce planning and internal mobility programs.
"NYC" becomes a structured object with city: "New York", region: "New York", country: "United States", ISO code: "US". "London, UK" resolves to city, region, country and ISO. The location normalization API transforms free-text location strings into structured geographic data. 100% resolution with 94.5% deterministic accuracy.
"Cambridge" alone could be Massachusetts or England. "Portland" could be Oregon or Maine. The location normalization engine uses context from the resume's other signals, including company locations, phone area codes and language patterns, to resolve ambiguous location references to the correct geographic entity.
Modern resumes increasingly list "Remote", "Hybrid", "WFH" or "Distributed" as location. The normalization API recognizes these work arrangement indicators, preserves them as structured tags and, where available, extracts the candidate's physical base location from other resume signals for geographic filtering.
The Resume Reader API is built to handle the full range of real-world document inputs, from cleanly formatted PDFs to photographed paper resumes. Each capability addresses a specific challenge in the document-to-data pipeline.
Accept resumes in PDF, DOCX, TXT, RTF, HTML, ODT, XLSX and XLS formats. The parser handles complex multi-column layouts, embedded tables, headers and footers, text boxes and nested formatting. Up to 10 MB per file and 20 pages per document. No format conversion required on the client side.
Supported formatsScanned paper resumes, photographed documents and image-based PDF files are automatically detected and processed through the optical character recognition pipeline. The OCR engine handles varying scan quality, rotation, skew and mixed text-and-image layouts. No preprocessing or separate OCR step required.
OCR capabilitiesSet the anonymize parameter to true and the API strips all personally identifiable information from the output: names, email addresses, phone numbers, physical addresses, photographs and social media profiles. The anonymized profile preserves skills, experience, education and qualifications for bias-free screening and GDPR-compliant data handling.
Anonymization guideThe API automatically computes total years of experience by calculating the duration of each work history entry and summing them with overlap detection. Per-role duration is computed individually. A seniority classification algorithm assigns each candidate to a level: intern, junior, mid-level, senior, lead, principal, director or executive.
The Resume Reader API never fabricates data. If a field is not present in the source document, the output returns null for that field rather than guessing or hallucinating content. This evidence-faithful design ensures that downstream systems can trust the structured output and distinguish between genuine candidate data and missing information.
Parse resumes written in any major language with full Unicode support. The API handles Latin, Cyrillic, CJK, Arabic and Devanagari scripts. Multilingual resumes with mixed-language content, such as a German resume listing English-named technologies and French certifications, are processed without language configuration.
From recruitment agencies processing thousands of resumes per week to HR tech products embedding parsing into their platforms, the Resume Reader API serves teams that need reliable, structured data from unstructured documents.
Staffing firms and executive search consultancies process hundreds of resumes per role. The API converts every submission into a searchable, filterable structured profile in seconds. 114+ fields per candidate enable deep matching against job requirements, client preferences and historical placement patterns. Seniority computation and experience calculation eliminate manual review of chronological timelines.
Applicant tracking systems receive resumes as file attachments. The Resume Reader API extracts the structured data that powers search, filtering and ranking inside the ATS. Skills arrays, experience objects and education records flow directly into database fields. Normalization APIs ensure consistent data across candidates who describe the same qualifications differently.
HR technology vendors embed the Resume Reader API into their products to offer resume parsing as a native feature. REST API integration with JSON request and response formats. 30 requests per 60 seconds rate limit. From $0.0198 per parse at volume. White-label ready with no end-user-facing branding. SDKs and documentation for fast integration.
Temporary staffing agencies manage high-volume, fast-turnaround candidate pipelines. The API processes resumes in bulk, extracting skills, availability, location and experience data for rapid matching against open assignments. Location normalization ensures geographic filtering works consistently regardless of how candidates describe their area.
Job boards and freelance platforms need structured candidate profiles that users can search, filter and compare. The API converts uploaded resumes into rich, searchable profiles without requiring candidates to fill out lengthy forms. Skills, experience and education data populate marketplace listings automatically.
Organizations implementing blind hiring workflows use the built-in anonymization feature to strip names, addresses, ages, gender indicators and photographs from candidate profiles. The anonymized output preserves every qualification, skill, experience entry and education record, enabling evaluation based solely on professional merit.
The Resume Reader API handles the full document processing pipeline in a single API call. Upload any supported file format and receive 114+ structured fields with optional normalization and anonymization.
The Resume Reader API handles structured data extraction from resumes and CVs. The Anonymization API extends document intelligence to broader redaction workflows across PDFs, DOCX, PPTX, XLSX and other enterprise document formats with 70+ PII type detection in 50+ languages.
The Anonymization API extends document intelligence beyond resumes. Automatically detect and redact PII, PHI and sensitive data from PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON and XML files. 70+ PII types detected with 99.9% accuracy across 50+ languages. Compliant with GDPR, CCPA and HIPAA requirements.
Anonymization APIProcess images and video files for automatic face detection, license plate blurring and screen content redaction. The computer vision pipeline handles photographs, scanned documents, video recordings and surveillance footage. Useful for document intelligence workflows involving identity verification documents, passport scans and medical imaging.
Image and video redactionTranscribe audio recordings with automatic PII removal. The speech-to-text pipeline detects and redacts names, account numbers, social security numbers and other sensitive information spoken during interviews, call recordings and meeting transcriptions. Output includes clean text and timestamped redaction metadata.
Audio redactionEnterprise reliability with transparent pricing. Every metric independently verified through 64 quality checks.
Send us a sample resume in any format. We will return the full 114-field structured JSON profile with normalization and anonymization applied, so you can evaluate parsing accuracy and field coverage before integrating. From $0.0198 per parse at production volume with plans starting at $99 per month.