Alpha Quantum ALPHA QUANTUM
Home About Contact
Solutions
E-Commerce Financial Services Healthcare Digital Marketing Legal & Compliance Content Moderation Data Privacy Customer Intelligence Document Intelligence Brand Safety
Industries
Healthcare Finance Retail Manufacturing Telecommunications Government Insurance Media Energy Education
Try Demo
Enterprise software since 2007 · 7 AI data platforms · 300+ organizations

7 AI data platforms
for the agent era

Alpha Quantum classifies the web at full scale and turns it into decisions machines can act on: what a page is, which AI tools employees may use and what they do with the data, and which pages an AI agent may open. Our AI block list covers 20,000+ AI-tool domains with risk flags on every row. Our AI agent allow list covers 40M domains with 28 page types each.

0
Domains classified
0
Agent allow list
0
AI-tool domains
0
Organizations
2026 AI agent incidents replayed against our 40M database + egress rules + host list1 / 9
all 9 replayed · 9 would have been blockedfull analysis →
Our platforms

Seven platforms, each built for its job

Domain intelligence, AI tool control, agent browsing policy, product classification, privacy and talent data. Every number below comes from the platform's own page.

Website Categorization API

Classify any URL in real time across IAB v2, IAB v3, IPTC and web-filtering taxonomies. 15+ data fields per request and offline databases from 10M to 120M domains.

Domains
120M
Categories
700+
Fields / request
15+
Explore platform
RISK FLAGGED

AI block list

20,183+ AI-tool domains in 18 categories and 165 subcategories, updated daily. Every row carries four risk flags: trains on your data, data sovereignty, abusive purpose, AI-native or AI-enabled.

AI-tool domains
20K+
Subcategories
165
Delivery formats
7
AI block list
PER-URL POLICY

AI agent allow list

For 40M+ domains, the verified URLs of 28 page types, login, signup, checkout, pricing and 24 more, with 700+ IAB and 59 filtering categories per domain. The policy data layer for agent guardrails.

Domains
40M+
Page types
28
Categories / domain
700+
AI agent allow list

Product Categorization API

AI product classification for Google Shopping, Shopify, Amazon and IAB taxonomies, with gender categorization, buyer persona mapping and custom classifiers. 5,574 Google categories, 200+ languages.

Google categories
5,574
Languages
200+
Taxonomies
Google, Shopify, Amazon, IAB
Explore platform

URL Categorization Database

120M+ domains classified by IAB and web-filtering taxonomies, enriched with company and firmographic data and web technology stacks. Offline files for firewalls, DNS filters, proxies, SIEMs and data warehouses.

Domains
120M+
Taxonomies
IAB + WF
Enrichment
firmo + tech
Explore platform

Anonymization API

Discover and anonymize PII, PHI and other sensitive data in text, documents, images and video. Multi-technique engine, context-aware detection, real-time streaming for live data, built for GDPR.

Sensitive data types
70+
Languages
50+
Inputs
text, docs, image, video
Explore platform

Resume Reader API

Parse any resume into clean structured JSON: 114+ candidate fields from PDF, DOCX, TXT, RTF, HTML, ODT, Excel and scanned documents via OCR. Normalizes titles, skills and locations.

Structured fields
114+
File types
8+
Pages per resume
20
Explore platform
The agent era

Two datasets for a web that machines now use

Employees paste company data into AI tools. AI agents open pages nobody reviewed. Both problems are lists of domains and URLs, with the right flags, kept current. That is what we build.

shadow AI prevention

AI block list

Blocking AI by name does not work. The row for openai.com says the API does not train on your data. The row for chatgpt.com says the consumer app does. That difference is the product: four risk flags on every one of 20,183+ domains, so policy follows risk, not brand. Real rows from the free sample:

domaincategorytrains on your datasovereigntyabuse
openai.comModels & Infrastructurenolownone
chatgpt.comGeneral assistantsyeslownone
claude.aiGeneral assistantsyeslownone
deepseek.comFoundation modelsyeshighnone
grammarly.comGrammar, AI-enabledyeslownone
notion.soOffice copilots, AI-enablednolownone
character.aiAI companionsyeslownone
crushon.aiAI companionsyesunknownnsfw
deeplivecam.netFace swapunknownunknowndeepfake
Copied verbatim from the free 50-row sample. Downloads carry the full multi-label category list per domain.
Trains on your dataThe flag that decides whether a paste becomes training data. Yes, no or unknown per domain.
Data sovereigntyWhere the data goes and under whose law. Low, medium, high or unknown.
Abusive purposeNudify, deepfake, uncensored and NSFW tools flagged so they can be blocked outright.
AI-native or AI-enabledA chatbot versus an office tool that grew an AI feature. Different policy, same list.
for web-browsing agents

AI agent allow list

A domain tells an agent nothing about the page. The same site holds documentation it should read and a checkout it must never touch. Every domain in this database carries the verified URLs of 28 page types, so your gateway checks the page type before the request is sent, and unknown sites are blocked by default.

40M+domains
28page types
700+categories per domain
GET /docs/apidocumentation · allow
GET /pricingpricing · allow
GET /account/loginlogin · deny
POST /cart/checkoutcheckout · deny
POST /wiki/editpost · deny
GET unknown-host.iounclassified · block

Built at the end of 2025. Replayed request by request against all nine 2026 agent incidents, our database plus egress rules would have blocked almost all of them. Refreshed quarterly so the policy stays true as the web moves.

The AI tool landscape, live from the AI block list

18 categories · 20,183 domains · domains per category, multi-label · updated daily
Our pipeline, in one picture

How our AI agent allow list, egress rules and host list stop each request

Our 40M-domain database, our egress rules and our host list, checked by your gateway before any request leaves. This is the mechanism behind every result in the incident replay above.

Your agentresearch, sales, shopping, browser automation
GET /docs/apiGET /account/loginPOST /cart/checkoutGET /wiki.pl?action=editGET 169.254.169.254GET unknown-host.io
Your gateway checks our database, our egress rules and our host list, in this order
1 · Our host listHigh-value and dangerous hosts by name: cloud metadata addresses, registries, model hubs. Deny or flag on the host alone.
host
2 · Our 40M-domain databaseThe page type of this exact URL on this domain: login, checkout, upload, post, docs, pricing. 28 types, every URL verified on the live site.
page type
3 · Our egress rulesURL patterns that mean a write, on any domain, whatever the method: wiki edits, WebDAV, plugin installs, API keys.
pattern
4 · Our default-deny ruleNot in our database, not in our egress rules, not on our host list? The request does not leave. Unknown sites are blocked until we know them.
unknown
Allowdocumentation, pricing, product, article pages
Loggrey areas your policy wants to see
Denylogin, checkout, upload, post, edit, keys, metadata, unknown hosts
the web
Why an escaped agent gets nowhere: every useful move after an escape is a request to a page type in our database, a URL pattern in our egress rules, or a host on our host list. The metadata address dies at our host list. The wiki edit dies at our rules. The login form dies at our database. The unknown app host dies at our default-deny rule. That is the whole mechanism behind the nine replays in the hero.
How it works

Four engines behind seven platforms

These are separate systems with separate data. One classifies the web, one classifies products, one finds and removes sensitive data, one reads documents. Each platform runs on the engine built for it.

Web classification engine

120M domains · 700+ categories · 28 page types

The domain corpus behind website categorization, the offline database, the AI block list, which is one slice of that 120M-domain filtering database, and the AI agent allow list, which adds verified page-type URLs on 40M of those domains.

Product classification engine

5,574 Google categories · 200+ languages

Models trained on product titles and descriptions, not web pages. Maps free-text listings to Google Shopping, Shopify, Amazon and IAB taxonomies, assigns buyer personas and gender, and trains custom classifiers on your own taxonomy.

Privacy engine

70+ sensitive data types · 50+ languages

Context-aware detection of PII and PHI in text, documents, images and video, then anonymization, pseudonymization or redaction with a multi-technique engine, including real-time streaming for live data.

Document engine

114+ fields · 8+ file types · OCR

Turns resumes in PDF, DOCX, TXT, RTF, HTML, ODT, Excel and scanned images into structured candidate JSON, normalizing titles, skills and locations for recruiting and HR workflows.

Problems we solve

Six problems, one control surface per problem

Each card is a decision a machine has to make millions of times a day, and the data that makes the decision defensible.

An agent about to open the wrong page

A browsing agent cannot tell a checkout from a product page until it has loaded it. By then the request is out. Our per-URL page types turn that into a rule the gateway checks before the click.

documentationpricingproductarticleloginsignupcheckoutuploadpostcommentpassword reset+ 17 more
40M+ domains with verified page-type URLs
replayed against 9 incidents of 2026: 9 would have been blocked
AI agent allow list

Company data pasted into AI tools nobody approved

The long tail is where the risk lives: 20,183+ AI-tool domains, most of them tools nobody has heard of, flagged for training on your data, data sovereignty and abusive purpose. Allow by policy, block by risk.

20,183+ AI-tool domains, 18 categories, daily updates
AI block list

An ad about to render next to the wrong story

Programmatic moves too fast for manual review. 700+ categories separate market analysis from fraud reports, and real-time classification keeps inclusion and exclusion lists current.

<100ms per URL, fits bid-request timing
Website Categorization API

A record that cannot be shared as it is

Patient files, KYC documents and contracts carry personal data in every clause. 70+ sensitive data types found and transformed, in 50+ languages.

Anonymization API

A domain the filter has never seen

CIPA, acceptable use and network security all rest on knowing what every domain is. 120M+ classified domains in the formats firewalls, proxies and DNS filters already read.

URL Categorization Database

Ten thousand sellers, ten thousand category names

Listings land in the wrong place and search returns noise. Every product mapped to Google Shopping, Shopify, Amazon or your own taxonomy in one call.

Product Categorization API
Under the hood

One request to the web classification engine

Your application sends a URL. The response carries the domain's classification across every taxonomy in your subscription, the page type of that exact URL when you license the AI agent allow list, and audience and quality data, from the real-time API or from a local copy you keep synced.

One site can be several things at once, so the engine returns multiple labels rather than a single answer. A streaming platform can be Video, User-Generated and Entertainment at the same time, and its login page is still a login page.

  • Response in under 100ms, or zero-latency lookups from an offline database
  • Multi-label classification so your logic can weigh every applicable category
  • Page type per URL for agent gateways: allow, log or deny before the request leaves
  • Category names built for readable logs and audit-ready reporting

Real-time classification

Requestexample-media.com/account/login
ClassifyIAB: News · IPTC: Business · WF: Safe
Page typelogin · agent verdict: deny
EnrichAudience: Finance, 25-54 · Quality: A · Risk: Low
Response in <100ms, cached for the next request
Who it is for

Built for the teams that need data to act

Different teams, different platforms. Every role gets the one built for the problem it owns.

IT & Security
AI & Agent teams
AdTech & Media
Compliance & Legal
Commerce & HR

You defend the network and answer for every domain that gets through

Enterprise IT teams, school districts and managed security providers deploy our databases directly into firewalls, DNS resolvers and proxies. Daily-refreshed data in RPZ, EDL, PAC and hosts formats keeps policies current.

The same enforcement points now govern AI tools: the AI block list adds 20,183+ AI-tool domains with risk flags to the filters you already run, so shadow AI is a policy, not a hope.

Daily-refreshed databases in RPZ, EDL, PAC, hosts, CSV and JSON formats
57+ web-filtering categories covering all CIPA-required content types
20,183+ AI-tool domains, four risk flags per row, 18 categories
Offline deployment for air-gapped networks and on-premise infrastructure

You ship or run agents that browse the open web

Research agents, sales agents, shopping agents and browser automation all land on pages nobody reviewed. The AI agent allow list gives your gateway a per-URL page type for 40M+ domains, so login, checkout, upload and post pages are refused before the request leaves.

Licensed as a database for embedding in your product, fully on-prem, or as a lookup API.

40M+ domains, verified URLs for 28 page types each
700+ IAB and 59 filtering categories on every domain for policy at the right resolution
Built end of 2025; replayed against the 2026 incidents, would have blocked almost all
Quarterly refreshes; database licence for embedding, or lookup API

You need brand safety and contextual signals at bid-request speed

SSPs, DSPs and publishers use real-time website categorization across 700+ categories to keep brand-safety inclusion and exclusion lists current against a web that changes every hour, protecting advertisers without killing reach.

IAB v2 and v3 taxonomies are supported natively, and sub-100ms responses fit programmatic timing.

700+ categories for granular brand-safety inclusion and exclusion lists
Sub-100ms API response fits programmatic bid-request timing
IAB v2 and v3 taxonomies supported natively
Offline databases for pre-bid filtering at scale

You navigate regulations that demand automated data handling

Healthcare systems, financial institutions and legal teams integrate the Anonymization API into document pipelines. Medical records are de-identified for research, financial narratives redacted for filings, contract reviews stripped of counterparty PII before storage.

Volumes that would need a 50-person team are handled by one endpoint, consistently and auditably.

Detection of 70+ sensitive data types in text, documents, images and video
Anonymization, pseudonymization and redaction in one engine
Real-time streaming for live data, built for GDPR
50+ languages for global document pipelines

You turn messy data into structured, actionable records

Marketplaces and e-commerce platforms use the Product Categorization API to normalize millions of seller-submitted listings into clean category trees, with buyer personas and gender labels alongside the category.

Recruiting and HR teams use the Resume Reader API to turn any resume file into 114+ structured fields, from $0.0198 per resume.

Product classification for Google Shopping, Shopify, Amazon and IAB taxonomies
5,574 Google categories, 200+ languages, custom classifier training
Resume parsing from 8+ file types including scanned documents via OCR
Normalized titles, skills and locations for clean candidate records
Why Alpha Quantum

What separates us from generic classification vendors

120M

Coverage generic vendors cannot match

Most providers cover the top 1 to 10 million domains. Our corpus covers 120 million, including the long tail of new sites, niche verticals and regional domains. If your use case depends on classifying the unusual, we cover it.

7

Purpose-built platforms, one vendor

Classification, AI tool control, agent browsing policy, product data, privacy and talent intelligence, each on the engine built for it. One vendor relationship, consistent APIs, one support channel.

14d

Test before you buy, always

Every API offers a 14-day free trial with no credit card. Database products include free sample downloads. Live demos run against production models with your own data, no registration walls.

Since 2007

Eighteen years of enterprise software

Before AI data platforms, Alpha Quantum built quantitative software for the financial industry: portfolio optimisation, risk management, security analysis and news analytics used by asset managers, hedge funds and wealth managers. The same engineering discipline now runs our data platforms.

quantitative finance

Portfolio Optimiser

Mean-variance, CVaR and multiperiod optimisation with efficient frontiers, correlation cleaning, backtesting over parameter grids and automated factsheet generation.

View product presentation (PDF)
quantitative finance

Risk Management

VaR and CVaR across historical, parametric and Monte Carlo methods, stress testing with historical and custom scenarios, risk attribution and pre-trade compliance.

View product presentation (PDF)
Try it yourself

Live demos and samples, no account required

Test the platforms with your own data. Demos run against production APIs. No sign-up, no credit card, no time limit.

Website Categorization

Enter any URL and see IAB categories, technology stack, audience personas and quality scores in real time.

Launch Demo

AI agent allow list

The 2026 agent incidents replayed request by request, with the page type or rule that refuses each one.

See the replay

AI block list

Download the free 50-row sample with categories and all four risk flags, and browse the 18 categories.

Get the sample

Product Categorization

Paste a product title and get instant classification across Google Shopping, Shopify and Amazon taxonomies.

Launch Demo

Redaction

Paste text with personal data and watch the API detect and redact names, emails, IDs and more.

Launch Demo

Anonymization

Anonymize and pseudonymize sensitive data in text and documents while keeping the data usable.

Launch Demo

Resume Reader

See a resume parsed into structured JSON with 114+ fields, from any common file type.

Launch Demo
Straight answers

Common assumptions about classification, addressed

01

“Generic classification tools are good enough”

Generic tools cover the popular web, the top 1 to 10 million domains. That leaves most of the internet unclassified, including the long-tail sites where brand-safety incidents happen and where agents wander. Our 120M-domain corpus closes that gap.
02

“We can block AI tools by name”

A few hundred well-known tools are easy. The other twenty thousand are where the risk lives, and even the famous ones differ: openai.com does not train on your data, chatgpt.com does. Risk flags per domain are what make the policy enforceable.
03

“A domain allow list is enough for agents”

A domain tells you nothing about the page. The same site holds documentation an agent should read and a checkout it must never touch. Per-URL page types are what make the rule expressible.
04

“Static databases go stale within weeks”

Ours do not. The web corpus takes in new and changed domains every day, the AI block list refreshes daily, the agent allow list quarterly, and offline subscribers receive exports on their schedule.
05

“Multi-vendor data stacks are too complex”

Agreed, which is why Alpha Quantum operates seven platforms under one roof. Security teams, agent builders, advertisers, compliance officers and data teams each get a purpose-built platform with one vendor relationship behind all of them.
06

“Classification accuracy does not move the needle”

It does when the downstream cost is real. A miscategorized domain blocks revenue or creates a PR crisis. A mis-labeled product tanks conversion. A missed PII entity triggers a fine. An agent that opens a checkout page spends money.
Use cases

How enterprises deploy Alpha Quantum platforms

Six deployments, each drawn as it runs: where the data comes from, which platform decides, and where the decision is enforced.

Shadow AISecurity teams

AI tool governance on the network

Employee trafficDNS, proxy, CASB logs
AI block list20,183+ domains, 4 risk flags
Firewall, DNS, proxyallow, log, block by flag
approved assistants stay ontrains on your data: loggedhigh sovereignty risk: blockednudify, deepfake: blocked
Formats: EDL, PAC, DNS RPZ, hosts, CSV, JSON, API · refreshed dailyAI block list
AI agentsAgent and platform teams

Browsing policy for agent fleets

Agent requestevery navigation, before it leaves
AI agent allow list40M domains, 28 page types
Agent gatewayallow, log, deny per URL
docs, pricing, product: allowedlogin, checkout, upload, post: deniedunknown host: blocked by default
Database licence for embedding, on-prem · lookup API · quarterly refreshAI agent allow list
SecurityIT, school districts, MSSPs

Network-level content filtering

Every domain requestedenterprise, campus, library
URL Categorization Database120M+ domains, IAB + filtering
Firewall, DNS resolver, proxypolicy by category
CIPA categories coveredacceptable use enforcedmalware, phishing categories blocked
RPZ, EDL, PAC, hosts, CSV, JSON · on-prem, air-gappedURL Categorization Database
AdTechSSPs, DSPs, publishers

Brand safety at bid-request speed

Bid requestpage URL, milliseconds to decide
Website Categorization API700+ categories, <100ms
Inclusion, exclusion listsupdated as the web changes
market analysis: suitablefraud reports: excludedIAB v2 and v3 native
Real-time API or offline database for pre-bid filteringWebsite Categorization API
ComplianceHealthcare, finance, legal

Automated sensitive-data handling

Records, filings, contractstext, documents, images, video
Anonymization API70+ sensitive data types, 50+ languages
Shared repository, research setanonymized, pseudonymized, redacted
de-identified for researchredacted for filingsreal-time streaming for live data
Built for GDPR, HIPAA and CCPA workflowsAnonymization API
Commerce and HRMarketplaces, recruiting teams

Catalog and candidate intelligence

Seller listings, resume filesfree text, any file type
Product Categorization + Resume Reader5,574 categories · 114+ fields
Clean category tree, ATS recordstructured, normalized
Google, Shopify, Amazon taxonomiesbuyer persona and gender labelstitles, skills, locations normalized
200+ languages · OCR for scanned resumesProduct Categorization
Questions

Common questions about Alpha Quantum

01What makes Alpha Quantum different from other classification providers?
Depth across seven distinct platforms on four separate engines. Our web corpus of 120M domains is the largest commercially available classification database, and every domain carries multiple taxonomies. On top of it we operate the AI block list and the AI agent allow list, and next to it separate engines for product categorization, anonymization and resume parsing.
02What is the AI block list and what are the risk flags?
A daily-updated list of 20,183+ AI-tool domains in 18 categories and 165 subcategories, delivered as EDL, PAC, DNS RPZ, hosts file, CSV, JSON and API. Every row carries four flags: whether the tool trains on your data, its data sovereignty risk, any abusive purpose such as nudify, deepfake, uncensored or NSFW, and whether it is AI-native or an AI-enabled product. Firewalls, DNS filters, proxies and CASBs use the flags to allow, log or block by policy. See the AI block list.
03What is the AI agent allow list?
A database of 40M+ domains where each domain carries the verified URLs of 28 page types, login, signup, checkout, pricing, upload, post, comment, documentation and more, plus 700+ IAB and 59 filtering categories. An agent gateway checks the page type before a request is sent and refuses writable or credential pages by rule. Built at the end of 2025 and replayed against the 2026 agent incidents, it would have blocked almost all of them. Refreshed quarterly. See the AI agent allow list.
04Why is a domain allow list not enough for agents?
A domain tells the agent nothing about the page. The same site holds documentation an agent should read and a checkout it must never touch. The 2026 incidents crossed exactly those lines: wiki edit endpoints, registry uploads, login forms. Per-URL page types are what make the rule expressible before the request leaves.
05Do I need an API, or can I use offline data?
Both. Real-time APIs deliver sub-100ms classification for live traffic. Offline databases run entirely inside your infrastructure, ideal for air-gapped networks, firewalls and environments where external calls are not permitted. Many customers use both: offline for bulk processing, API for new or unknown URLs.
06How fresh is the data?
The web corpus takes in new and changed domains every day and API responses reflect the latest classification. The AI block list refreshes daily. The AI agent allow list refreshes quarterly. Offline database subscribers receive exports on their license schedule.
07What industries do you serve?
Telecommunications, advertising technology, cybersecurity, AI and agent platforms, publishing, e-commerce, financial services, healthcare, education, government, energy, insurance, manufacturing and media. The common thread is a need to classify, filter, govern or enrich large volumes of URLs, products, documents and content.
08Is there a free trial?
Yes. Every API platform offers a 14-day free trial with no credit card. Database products include free sample downloads so you can validate format and quality before purchasing. Live demos are available without registration.
09How do you handle data privacy and security?
We classify publicly available web content and do not collect, store or process user-level browsing data. The Anonymization API is designed to help customers meet GDPR, HIPAA, CCPA and similar requirements. API traffic is encrypted in transit and query data is not retained beyond the processing window.
10Can I integrate with my existing infrastructure?
Yes. We deliver data in every format enterprise infrastructure consumes: REST API, CSV, JSON, SQLite, DNS RPZ, PAC files, EDL feeds and hosts files. Palo Alto, Squid, BIND, Snowflake or a custom application, the data fits without conversion scripts.
11How large is the product categorization coverage?
The Product Categorization API classifies products across 5,574 categories in the Google Shopping taxonomy, plus Shopify, Amazon and IAB taxonomies, supports 200+ languages, handles free-text product titles, and includes buyer persona mapping, gender categorization and custom classifier training.

Find the platform built for your problem

Start with a demo, download a sample database, or talk to our team about your use case. No commitment, no sales pressure, just data you can test.