Alpha Quantum ALPHA QUANTUM
Home
Platforms
Solutions
Industries
About Contact
Try Demos
Document Intelligence

Extract structured data from any document with AI-powered parsing, 114-field extraction and built-in anonymization

One API call converts resumes in PDF, DOCX, scanned images and five other formats into 114+ structured JSON fields. Three normalization APIs standardize job titles, skills and locations. Built-in anonymization removes candidate identity for bias-free screening. From $0.0198 per parse with multilingual OCR support.

Resume Reader API

114+ fields per resume

Convert any resume file into a structured JSON candidate profile. AI-powered parsing with OCR, seniority computation, skill normalization and built-in anonymization for bias-free screening workflows.

Fields
114+
File Types
8+
Norm. APIs
3
Cost
$0.0198
Explore platform
The Challenge

Unstructured documents block every hiring pipeline

Resumes arrive in dozens of formats, use inconsistent terminology and contain sensitive personal information that must be handled carefully. Manual processing is slow, expensive and introduces errors and bias at every step of the recruitment workflow.

Resumes arrive in every format imaginable

PDF documents with complex layouts. DOCX files with embedded tables and columns. Scanned paper resumes saved as image-based PDFs. Rich text files from legacy systems. HTML exports from LinkedIn and job boards. Even XLSX spreadsheets from staffing agencies. The Resume Reader API accepts PDF, DOCX, TXT, RTF, HTML, ODT, XLSX and XLS files up to 10 MB and 20 pages each, with built-in OCR for scanned and image-based documents.

Manual data entry is expensive and error-prone

Recruiters spend hours copying candidate information from documents into applicant tracking systems. Every manual keystroke introduces the possibility of transposition errors, missed fields and inconsistent formatting. At scale, a team processing hundreds of resumes per day loses significant time to data entry that could be spent on candidate evaluation and relationship building.

Normalization chaos across job titles and skills

The same role appears as "Sr. SWE II", "Senior Software Engineer", "Staff Developer" and "Lead Programmer" across different resumes. Skills appear as "K8s", "Kubernetes", "k8s orchestration" and "container orchestration". Locations read as "NYC", "New York City", "New York, NY" and "Manhattan". Without normalization, matching, filtering and reporting break down. The three normalization APIs resolve 100% of inputs with deterministic accuracy rates of 95.4%, 98.2% and 94.5% respectively.

Bias in screening requires anonymization capability

Research consistently shows that candidate names, ages, gender indicators and addresses influence hiring decisions. Organizations pursuing equitable hiring practices need a reliable way to remove identifying information from resumes before human review. The Resume Reader API includes a built-in anonymization parameter that strips all personal identifiers while preserving the skills, experience and qualifications that matter for evaluation.

Structured Output

114+ fields extracted from every resume

The Resume Reader API converts unstructured document content into a comprehensive structured JSON profile. Every field is evidence-faithful: if the information is not present in the source document, the API returns null rather than fabricating data. These are the core field groups returned per candidate.

Field GroupTypeExample OutputUse Case
Personal Informationobjectname, email, phone, location, LinkedInCandidate profile creation
Work Experiencearraytitle, company, start/end dates, descriptionCareer history reconstruction
Educationarraydegree, institution, graduation year, GPAQualification verification
SkillsarrayPython, Kubernetes, SQL, LeadershipSkill matching and filtering
CertificationsarrayAWS Solutions Architect, PMP, CPACredential verification
LanguagesarrayEnglish (native), German (fluent), French (B2)Global and multilingual hiring
Seniority LevelcomputedSenior / Lead / Executive / Entry-LevelRole-level filtering and routing
Total Experiencecomputed12.3 years (with per-role durations)Experience-based screening
Publicationsarraytitle, journal/conference, year, co-authorsAcademic and research roles
Projectsarrayname, description, technologies, outcomesPortfolio and capability assessment
Full field reference and JSON schema on resumereaderapi.com. All 114+ fields returned per parse. 64 out of 64 quality checks passed in independent evaluation.
resume.pdf114 Fields · Work History · Skills · Education
candidate.docxOCR · Layout Analysis · Table Extraction
scanned_cv.pdfImage OCR · Text Reconstruction · Field Mapping
bulk_upload.xlsxBatch Parse · Normalized JSON · ATS Import
Normalization APIs

Three normalization APIs that standardize every data point

Raw resume data is inconsistent by nature. The same job title, skill or location appears in dozens of variations across candidates. The three normalization APIs resolve 100% of inputs into standardized forms at 0.1 credits per item, enabling consistent matching, filtering and reporting across your entire candidate database.

Job Title Normalization
Skill Normalization
Location Normalization

Canonical Title Mapping

"Sr. SWE II" becomes "Senior Software Engineer". "VP of Eng" becomes "Vice President of Engineering". "Dev Lead" becomes "Development Lead". The job title normalization API maps every variation, abbreviation and informal title to a canonical form. 100% of inputs resolved with 95.4% deterministic accuracy.

Seniority Classification

Each normalized title is tagged with a seniority level: intern, junior, mid-level, senior, lead, principal, director, vice president, C-suite. The classification enables experience-based routing and compensation benchmarking across your entire candidate pool.

Cross-Industry Mapping

Job titles vary dramatically across industries. A "Producer" in media is different from a "Producer" in manufacturing. The normalization engine uses context from the resume's industry signals and company information to disambiguate titles and assign the correct canonical form and seniority level.

Abbreviation Resolution

"K8s" becomes "Kubernetes". "Phyton" becomes "Python". "JS" becomes "JavaScript". "ML" becomes "Machine Learning". "TF" becomes "TensorFlow". The skill normalization API handles abbreviations, misspellings, acronyms and informal names. 100% resolution rate with 98.2% deterministic accuracy across the full technology and business skill landscape.

Synonym Clustering

"React.js", "ReactJS", "React" and "React Framework" all map to a single canonical "React" entry. "PostgreSQL", "Postgres" and "PG" converge to "PostgreSQL". Synonym clustering eliminates duplicate skill entries and enables accurate skill frequency analysis across candidate pools.

Category Tagging

Each normalized skill is tagged with its category: programming language, framework, database, cloud platform, methodology, soft skill and dozens more. Category tags enable structured skill inventory reports and gap analysis for workforce planning and internal mobility programs.

Structured Geography

"NYC" becomes a structured object with city: "New York", region: "New York", country: "United States", ISO code: "US". "London, UK" resolves to city, region, country and ISO. The location normalization API transforms free-text location strings into structured geographic data. 100% resolution with 94.5% deterministic accuracy.

Ambiguity Resolution

"Cambridge" alone could be Massachusetts or England. "Portland" could be Oregon or Maine. The location normalization engine uses context from the resume's other signals, including company locations, phone area codes and language patterns, to resolve ambiguous location references to the correct geographic entity.

Remote and Hybrid Handling

Modern resumes increasingly list "Remote", "Hybrid", "WFH" or "Distributed" as location. The normalization API recognizes these work arrangement indicators, preserves them as structured tags and, where available, extracts the candidate's physical base location from other resume signals for geographic filtering.

Full normalization API documentation
Processing Capabilities

Six capabilities that handle every document scenario

The Resume Reader API is built to handle the full range of real-world document inputs, from cleanly formatted PDFs to photographed paper resumes. Each capability addresses a specific challenge in the document-to-data pipeline.

Multi-Format Parsing

Accept resumes in PDF, DOCX, TXT, RTF, HTML, ODT, XLSX and XLS formats. The parser handles complex multi-column layouts, embedded tables, headers and footers, text boxes and nested formatting. Up to 10 MB per file and 20 pages per document. No format conversion required on the client side.

Supported formats

OCR for Scanned Documents

Scanned paper resumes, photographed documents and image-based PDF files are automatically detected and processed through the optical character recognition pipeline. The OCR engine handles varying scan quality, rotation, skew and mixed text-and-image layouts. No preprocessing or separate OCR step required.

OCR capabilities

Built-In Anonymization

Set the anonymize parameter to true and the API strips all personally identifiable information from the output: names, email addresses, phone numbers, physical addresses, photographs and social media profiles. The anonymized profile preserves skills, experience, education and qualifications for bias-free screening and GDPR-compliant data handling.

Anonymization guide

Seniority Computation

The API automatically computes total years of experience by calculating the duration of each work history entry and summing them with overlap detection. Per-role duration is computed individually. A seniority classification algorithm assigns each candidate to a level: intern, junior, mid-level, senior, lead, principal, director or executive.

Evidence-Faithful Output

The Resume Reader API never fabricates data. If a field is not present in the source document, the output returns null for that field rather than guessing or hallucinating content. This evidence-faithful design ensures that downstream systems can trust the structured output and distinguish between genuine candidate data and missing information.

Multilingual Support

Parse resumes written in any major language with full Unicode support. The API handles Latin, Cyrillic, CJK, Arabic and Devanagari scripts. Multilingual resumes with mixed-language content, such as a German resume listing English-named technologies and French certifications, are processed without language configuration.

Use Cases

How teams use document intelligence

From recruitment agencies processing thousands of resumes per week to HR tech products embedding parsing into their platforms, the Resume Reader API serves teams that need reliable, structured data from unstructured documents.

01

Recruitment Agencies and Executive Search

Staffing firms and executive search consultancies process hundreds of resumes per role. The API converts every submission into a searchable, filterable structured profile in seconds. 114+ fields per candidate enable deep matching against job requirements, client preferences and historical placement patterns. Seniority computation and experience calculation eliminate manual review of chronological timelines.

02

ATS Enrichment and Candidate Databases

Applicant tracking systems receive resumes as file attachments. The Resume Reader API extracts the structured data that powers search, filtering and ranking inside the ATS. Skills arrays, experience objects and education records flow directly into database fields. Normalization APIs ensure consistent data across candidates who describe the same qualifications differently.

03

HR Tech Product Integration

HR technology vendors embed the Resume Reader API into their products to offer resume parsing as a native feature. REST API integration with JSON request and response formats. 30 requests per 60 seconds rate limit. From $0.0198 per parse at volume. White-label ready with no end-user-facing branding. SDKs and documentation for fast integration.

04

Staffing and Temp Agencies

Temporary staffing agencies manage high-volume, fast-turnaround candidate pipelines. The API processes resumes in bulk, extracting skills, availability, location and experience data for rapid matching against open assignments. Location normalization ensures geographic filtering works consistently regardless of how candidates describe their area.

05

Job Boards and Talent Marketplaces

Job boards and freelance platforms need structured candidate profiles that users can search, filter and compare. The API converts uploaded resumes into rich, searchable profiles without requiring candidates to fill out lengthy forms. Skills, experience and education data populate marketplace listings automatically.

06

CV Anonymization for Bias-Free Hiring

Organizations implementing blind hiring workflows use the built-in anonymization feature to strip names, addresses, ages, gender indicators and photographs from candidate profiles. The anonymized output preserves every qualification, skill, experience entry and education record, enabling evaluation based solely on professional merit.

Processing Pipeline

From uploaded file to structured JSON in seconds

The Resume Reader API handles the full document processing pipeline in a single API call. Upload any supported file format and receive 114+ structured fields with optional normalization and anonymization.

01
Upload
PDF, DOCX, TXT, RTF, HTML, ODT, XLSX or XLS. Up to 10 MB and 20 pages per file.
02
Parse
AI extraction of 114+ fields. OCR for scanned documents. Layout and table analysis.
03
Normalize
Job titles, skills and locations standardized via three normalization APIs at 0.1 credits each.
04
Deliver
Structured JSON output. Anonymized if requested. Ready for ATS, database or API integration.
Platforms

Two platforms powering document intelligence

The Resume Reader API handles structured data extraction from resumes and CVs. The Anonymization API extends document intelligence to broader redaction workflows across PDFs, DOCX, PPTX, XLSX and other enterprise document formats with 70+ PII type detection in 50+ languages.

Document Redaction

The Anonymization API extends document intelligence beyond resumes. Automatically detect and redact PII, PHI and sensitive data from PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON and XML files. 70+ PII types detected with 99.9% accuracy across 50+ languages. Compliant with GDPR, CCPA and HIPAA requirements.

Anonymization API

Visual Content Redaction

Process images and video files for automatic face detection, license plate blurring and screen content redaction. The computer vision pipeline handles photographs, scanned documents, video recordings and surveillance footage. Useful for document intelligence workflows involving identity verification documents, passport scans and medical imaging.

Image and video redaction

Audio Transcription Redaction

Transcribe audio recordings with automatic PII removal. The speech-to-text pipeline detects and redacts names, account numbers, social security numbers and other sensitive information spoken during interviews, call recordings and meeting transcriptions. Output includes clean text and timestamped redaction metadata.

Audio redaction
Scale

Production-grade document processing infrastructure

Enterprise reliability with transparent pricing. Every metric independently verified through 64 quality checks.

114+
Fields Extracted
8+
File Types Supported
3
Normalization APIs
$0.0198
Per Parse
64/64
Quality Checks Passed
30/60s
Rate Limit
10 MB
Max File Size
20
Max Pages Per File

Parse your first resume for free

Send us a sample resume in any format. We will return the full 114-field structured JSON profile with normalization and anonymization applied, so you can evaluate parsing accuracy and field coverage before integrating. From $0.0198 per parse at production volume with plans starting at $99 per month.

Contact Us Resume Reader API