Engagement Foundation Review

Tonic.ai
Audit Foundation

AI assistants are becoming the first place teams ask how to strip PII, PHI, and PCI out of documents and transcripts before that data reaches a model — and the vendors named in those answers today are locking in an advantage while the category is still being defined. Before we run the audit, we need to make sure we're asking the right questions about the right competitors to the right buyers. This document presents what we've learned about Tonic.ai's market — your job is to tell us what we got right, what we got wrong, and what we missed.

Prepared July 27, 2026
tonic.ai
Unstructured data de-identification & redaction
GEO Readiness

Where You Stand Today

Before we measure how often Tonic.ai is cited in unstructured de-identification and redaction questions, these three signals tell us whether AI crawlers can reach the site, read it, and tell what is current. All three are derived mechanically from the Layer 1 crawl.

Technical Readiness
Needs Attention
Four high-severity findings, zero critical. Most consequential: all three /alternatives/ pages render an orphaned block reading "Last updated December 2023. Comparison based on Delphix's full suite of services" — including the Limina page, the head-to-head against the closest competitor in this graph. Also high: 1,878 sitemap URLs with zero <lastmod>, and the Synthesized and Delphix comparisons each published at multiple self-canonical URLs.
Content Freshness
Needs Attention
Weighted freshness: 0.61. Of 24 dated content-marketing pages, 12 were updated within 90 days, 8 are older than 6 months, and 3 are older than a year. 23 product and commercial pages carry no detectable date — verify manually. No sitemap URL publishes a <lastmod>, so nothing corroborates page age externally. For scale: AI-cited content runs 25.7% fresher on average than content in Google organic results (Ahrefs, August 2025).
Crawl Coverage
Needs Attention
robots.txt permits all seven declared AI crawlers — GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Googlebot, Bytespider — disallowing only /styleguide, and docs.tonic.ai publishes Content-Signal: ai-train=yes. Crawler access is still unverified: every response carries server: cloudflare, and edge bot rules operate above robots.txt. Sitemap is one flat file of 1,878 URLs, 1,443 of them (76.8%) per-release changelog entries.
Executive Summary

What You Need to Know

Unstructured data de-identification is a category being defined right now, and a large share of the buyers defining it start by asking an assistant rather than a search engine — 94% of B2B buyers now use LLMs somewhere in the buying process (6sense, November 2025). That makes the answer to "how do I de-identify documents for LLM training" a contested asset, and the vendors named in it today accumulate an advantage that compounds, because brand web mentions predict AI citation far more strongly than backlinks or organic traffic do (Seer Interactive, October 2025) — visibility feeds the signal that produces more visibility. Tonic.ai enters this with unusual raw material for the category — a published, reproducible accuracy benchmark and a self-hosted deployment story — but raw material only converts into citations if AI systems can reach it, parse it, and tell that it is current.

This Foundation Review is the input layer for that measurement, and it asks you to confirm or correct four things. The competitive landscape determines which vendors we test Tonic.ai against head-to-head and which we merely watch. The buyer personas determine the intent, vocabulary, and altitude of every query we generate. The feature and pain-point taxonomies supply the actual language buyers use when they ask. And the Layer 1 technical baseline assesses whether AI systems can access and interpret tonic.ai at all, independent of what the content says. Nothing here is a finished conclusion — it is the set of assumptions the audit gets built on, and an assumption is far cheaper to fix now than a query set is to re-run later.

The validation call is a working session, not a presentation, and two kinds of decisions come out of it. The first is input validation: are the right entities in the right tiers, do the personas match who actually shows up in your deals, and are the capability ratings honest relative to the vendors you meet in live evaluations? Those answers set the audit architecture — query volume, competitive matchups, and the vocabulary every prompt is written in, across the AI platforms selected for this engagement. The second is engineering triage: several Layer 1 findings need no decision from anyone and can be in flight before we meet, so the audit measures an improved baseline rather than a known-broken one. The Pre-Call Checklist collects both lists in one place.

TL;DR — Action Items
  • 🟡 High: robots.txt allows every AI crawler — confirm Cloudflare is not blocking them above it — Engineering should check Security > Bots for Bot Fight Mode and the "Block AI bots" managed rule, then confirm from logs that GPTBot, ClaudeBot, PerplexityBot, and Google-Extended actually received 200s over the last 30 days.
  • 🟡 High: Competitor comparison pages carry stale, contradictory, and wrong-competitor "Last updated" stamps — Remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template and bind the visible date to the same CMS field that feeds JSON-LD dateModified.
  • 🟣 Validate at the Call: Real-Time Prompt & Response Redaction for Live Agents — We rated this weak because every path we could verify is batch or request/response rather than an inline low-latency proxy; if Textual has a production runtime path with published latency, we generate head-to-head queries against Skyflow, Protecto, and Nightfall instead of treating runtime as a conceded category.
  • 🟣 Validate at the Call: Terrence Whitlock — Manager, Records & Information Governance — He is the only persona we inferred rather than sourced, and Guided Redaction is documented as beta; if FOIA and litigation redaction is not yet a funded buying motion, his entire query vocabulary gets reallocated to the AI platform and security personas.
  • ✅ Start Now: emit <lastmod> across sitemap.xml and collapse the duplicate comparison URLs — Neither waits on the call: the sitemap fix is a Webflow CMS field binding, and the duplicate fix is choosing /vs/{competitor} as the canonical home and 301-ing /blog/tonic-vs-synthesized to it.
  • 📋 Validation Call: does this audit measure Tonic Textual, or the full Tonic.ai platform? — We scoped this graph to unstructured de-identification from the Textual seed URL; a platform-wide scope pulls Delphix, K2View, and Synthesized into the competitive set and changes the personas, features, and query vocabulary with it.
Orientation

How This Works

Three things worth knowing before you read the rest of this document.

What This Is This is the knowledge graph the audit will run against. Everything downstream — the buyer queries we generate, the competitors we test head-to-head, the language we phrase prompts in — is derived from the personas, competitors, capabilities, and pain points on these pages. It was built outside-in: from tonic.ai's own pages, competitor positioning, category listings, and review data, without access to your CRM or win/loss records. That is deliberate — it mirrors what an AI system can see about the unstructured de-identification and redaction category. It also means you know things we cannot.

What We Need From You Read the purple boxes. Each one names a specific entity we are uncertain about and states what changes in the audit if our reading is wrong. You do not need to review every card in detail — the purple questions are where your answer actually moves something. Everything else is here so you can check our work if you want to. The Pre-Call Checklist near the end collects every question in one printable list.

Confidence Badges Every entity carries a confidence badge. High means it came straight from a primary source — tonic.ai's own pages, a competitor's positioning, or a published review. Medium means we triangulated it from more than one indirect signal. Low means we inferred it and are asking you to confirm or kill it. Low-confidence items are not filler; they are the specific places where your judgment beats our research.

Company Profile

Who We Think You Are

How AI systems would describe Tonic.ai from the outside — the naming, category, and positioning every generated query is anchored to.

Client Profile

Company name Tonic.ai High
Domain tonic.ai
Name variants tracked Tonic · Tonic AI · TonicAI · Tonic.ai Inc. · Tonic Textual · Textual · Tonic Textual by Tonic.ai
Category Unstructured data de-identification and redaction — detecting and removing PII, PHI, and PCI from free text, documents, images, and audio, replacing it with realistic synthetic values so the data stays usable for LLM training, RAG, and analytics
Company segment Startup — ~95–105 employees, $45M raised through a 2021 Series B
Key products Tonic Textual · Tonic Structural · Tonic Fabricate
Positioning (from site) "Instantly redact sensitive PII from text and audio to safely train models, and secure real-time agentic workflows — the essential privacy layer for every LLM interaction."
Source Automated scrape — tonic.ai product, pricing, and solutions pages

→ Validate Is this audit measuring Tonic Textual, or the whole Tonic.ai platform? The seed URL was the Textual product page, so we scoped the category, the competitors, and the twelve capabilities to unstructured de-identification — which is why Delphix, K2View, and Synthesized are absent even though tonic.ai maintains dedicated /alternatives/ and /vs/ pages for all three. If the audit should cover Structural and Fabricate too, the test-data-management competitive set comes back in, the personas shift toward QA and platform engineering, and the query vocabulary changes from "redact PHI in clinical notes" to "database masking for test environments." One more wrinkle to settle in the same breath: the schema classifies Tonic.ai as a startup on headcount and funding, but the named customer list runs JPMorgan Chase, Qualcomm, Philips, Commonwealth Bank, and the NYC Department of Health — should query language read enterprise-evaluation even though the company is startup-segment?

Buyer Personas

Who Buys This

6 personas — 3 decision-makers, 1 evaluator, 2 influencers — each of whom searches the unstructured de-identification category differently, which is what determines the intent and vocabulary of every query we generate.

Critical Review Area This is the section where your correction is worth the most. Personas drive query generation directly: each one produces a distinct cluster of prompts written at their altitude, in their vocabulary, with their evaluation criteria. A persona who does not exist in your deals burns query budget on the wrong buyer. A persona who exists but is missing here means an entire buying conversation goes unmeasured.

Data Sourcing Note Role, department, seniority, influence level, veto power, and technical level come straight from the knowledge graph and carry the confidence badge shown on each card. The role description, primary buying jobs, and query focus areas are synthesized — our reading of what that role does during a de-identification evaluation, based on the capabilities and pain points they map to. Names are placeholders for the role, not real people. Correct the synthesized lines freely; they are our inference, not your data.

Priya Raghunathan
Head of AI Platform Engineering · Engineering · Director
Evaluator High
Owns the shared platform internal AI teams build on — model access, data pipelines, and the de-identification layer that feeds fine-tuning and RAG. Runs the technical bake-off and writes the recommendation that the budget holder signs.
Veto power: No — high influence over the choice, but the spend is approved above her
Technical level: High
Primary buying jobs: Shortlisting vendors, running accuracy bake-offs against internal corpora, proving the pipeline integrates with existing infrastructure without a rewrite
Query focus areas: Detection precision and recall on real-world text, self-hosted deployment, Python SDK and REST API design, throughput on large document archives
Source: Automated scrape — tonic.ai product, pipelines, and developer documentation

Priya drives the evaluation but holds no veto in our model — if AI platform leads at your accounts actually own the spend, we promote her to decision-maker and add procurement-stage queries (pricing per 1,000 words, TCO against an in-house Presidio build) to her cluster.

Marcus Delgado
Chief Information Security Officer · Information Security · C-Suite
Decision-maker High
Decides whether regulated text may be processed at all, and by whom. Owns the security review that eliminates vendors before technical evaluation starts — the reason "just use Comprehend" dies in the first meeting at some accounts and survives at others.
Veto power: Yes — can kill a vendor on trust-boundary grounds alone, regardless of accuracy
Technical level: Medium
Primary buying jobs: Vendor security review, data residency and trust-boundary approval, third-party risk sign-off, breach-exposure assessment
Query focus areas: Self-hosted and VPC deployment, SOC 2 and BAA posture, re-identification risk, whether regulated text ever leaves the environment
Source: Automated scrape — tonic.ai security, deployment, and compliance pages

We modeled Marcus as a hard veto on data leaving the VPC — if security instead signs off on a managed endpoint under a BAA, self-hosting stops being a gating query theme and Amazon Comprehend and Google Cloud DLP become live options inside his cluster rather than pre-eliminated ones.

Dana Okonkwo
Director of Privacy & Regulatory Compliance · Legal & Compliance · Director
Decision-maker High
Answers to auditors and regulators for whether "de-identified" data is genuinely de-identified. Translates HIPAA, GDPR, and CCPA obligations into what the data team is permitted to do, and carries the personal exposure when the answer is wrong.
Veto power: Yes — approval is a regulatory sign-off, not a preference
Technical level: Low — evaluates evidence and documentation, not architecture
Primary buying jobs: Regulatory defensibility review, expert-determination sign-off, defining what audit evidence must exist before data moves
Query focus areas: HIPAA expert determination, GDPR pseudonymization vs. anonymization, defensible audit trails, documentation an auditor will actually accept
Source: Automated scrape — tonic.ai expert-determination and compliance capability pages

Dana's query cluster assumes HIPAA expert determination is the deciding evidence — if GDPR and EU data residency is the more common blocker in your pipeline, we swap that set toward pseudonymization-under-GDPR language, which is a materially different vocabulary and surfaces a different competitive field.

Ethan Brandt
Director of Data Science & Machine Learning · Data & Analytics · Director
Influencer Medium
Consumes the de-identified corpus and feels over-redaction directly — his models degrade when semantic context is blacked out. Rarely signs the contract, but his verdict on output quality can end an evaluation.
Veto power: No — medium influence, exercised through evaluation results rather than authority
Technical level: High
Primary buying jobs: Judging output quality, benchmarking model performance on synthesized vs. raw data, sizing what gets unlocked if blocked datasets become usable
Query focus areas: Over-redaction and utility loss, synthetic replacement realism, multilingual and audio corpus coverage, chunking and extraction for vector databases
Source: Review mining — G2 reviewer titles and published case studies

Ethan is sourced from review mining at medium confidence and modeled as an influencer who inherits the evaluation — if data science actually initiates it, he becomes the entry-point persona and "our redaction script turned every document into a wall of [REDACTED]" moves to the top of the query set rather than the middle.

Sofia Halvorsen
VP of Engineering · Engineering · VP
Decision-maker Medium
Owns the platform budget line and the build-vs-buy call — specifically whether the team keeps funding two engineers to maintain an in-house Presidio pipeline or buys the problem away. Cares about consolidation as much as capability.
Veto power: Yes — assumed. The KG assigns her the budget line; that split was inferred, not observed.
Technical level: High
Primary buying jobs: Build-vs-buy decision, budget approval, consolidating structured and unstructured tooling under one contract and one security review
Query focus areas: Build vs. buy for PII redaction, total cost of ownership against open source, ongoing maintenance burden, vendor consolidation
Source: Review mining — G2 reviews and case studies; budget authority inferred

We gave Sofia veto power on the assumption VP Engineering owns the platform budget — if that money actually sits with the CISO or inside Priya's AI platform org, then one of our three veto-holders is wrong and the approval-stage queries are written for the wrong reader entirely.

Terrence Whitlock
Manager, Records & Information Governance · Legal Operations · Manager
Influencer Low
Runs manual redaction for FOIA responses, litigation production, and partner document sharing. Measures the problem in paralegal-weeks and in near-misses, and buys a workflow with a review screen rather than an API.
Veto power: No — medium influence, typically sponsored by legal or compliance leadership
Technical level: Low
Primary buying jobs: Replacing manual redaction labor, establishing a defensible review-and-sign-off process, getting batch throughput on production sets
Query focus areas: Automated document redaction software, FOIA and litigation production redaction, human review and approval workflows, scanned PDF and image handling
Source: LLM inference — Guided Redaction capability pages and the ilionx government redaction release; no review or customer data

Terrence is the only persona we inferred rather than sourced, and Guided Redaction is documented as beta — if records and legal redaction is not yet a funded buying motion, his entire query vocabulary gets reallocated to the AI platform and security personas rather than testing a market you are not selling into yet.

→ Who Else Shows Up? Three roles appear in de-identification deals often enough to be worth a check, and each would earn a dedicated query cluster if they show up in yours: a Chief Medical Information Officer or clinical data governance lead (if healthcare deals route through a clinical approval chain that is separate from the CISO's); a Data Protection Officer for EU entities (if GDPR is a distinct buying conversation from HIPAA rather than a subsection of it); and a third-party risk / vendor security reviewer (if the security questionnaire is its own gate with its own evaluation criteria, rather than something Marcus's team absorbs). Who else shows up in your deals?

Competitive Landscape

Who You're Measured Against

11 competitors identified — 6 primary and 5 secondary — spanning named privacy vendors, hyperscaler APIs, and one open-source framework.

Why Tiers Matter Tier assignment decides where query budget goes. Each primary competitor gets a dedicated head-to-head set — roughly six to eight prompts apiece, so about 36–48 of the audit's queries — testing direct comparisons like "Tonic Textual vs Microsoft Presidio," "best alternative to Private AI for PHI redaction," and "Comprehend vs Tonic for clinical note de-identification." Secondary competitors are measured only for co-occurrence: do they show up when a buyer asks an open category question. Two of the six primaries — Protecto and Skyflow — carry medium confidence because we tiered them on category overlap and search co-occurrence rather than observed deal data. One more thing worth flagging before queries are written: Private AI rebranded to Limina in March 2026, models trained before then will overwhelmingly say "Private AI," and your own comparison page still lives at /vs/privateai-vs-tonic — we collapse both names to one entity so head-to-head share against your closest rival is not split in half and undercounted.

Primary Competitors — 6

Limina

Primary High
getlimina.ai — formerly Private AI
The closest head-to-head rival — a stateless, container-based entity detection and redaction engine covering 50+ entity types across 52 languages in text, PDFs, images, audio, and DICOM, deployed entirely inside the customer's VPC. Strong on detection breadth and data residency, but unstructured-only with no structured or test-data companion product, and Tonic's own comparison page attacks it on the lack of a collaborative UI, custom entity tuning, and expert-determination partnerships. Rebranded from Private AI in March 2026, so AI responses use both names.
Source: Competitor site + tonic.ai comparison page

Microsoft Presidio

Primary High
microsoft.github.io/presidio — open source
The free open-source Python framework that every technical evaluator benchmarks against and many teams quietly ship instead of buying. Zero license cost and full in-house control, but a recognizer-and-regex architecture produces recall gaps on messy real-world text, and it ships with no UI, no synthesis, no audit trail, and no support — the buying team inherits an ongoing ML maintenance project.
Source: Category listings and developer community references

Amazon Comprehend

Primary High
aws.amazon.com/comprehend
Managed AWS NLP service with PII detection plus a Comprehend Medical variant for PHI — the "we already have AWS" default that costs nothing extra in procurement. Loses on accuracy in Tonic's October 2025 benchmark across legal, EHR, and call-transcript text, offers redaction but not context-preserving synthesis, and requires sending regulated text out to a managed endpoint.
Source: Competitor site + tonic.ai published benchmark

Google Cloud Sensitive Data Protection

Primary High
cloud.google.com — formerly Cloud DLP
Mature managed inspection and de-identification service spanning text, images, and structured data, with tokenization and format-preserving encryption built in — the strongest cloud-native option and a real evaluation contender for GCP shops. Weaker on free-text nuance (infoType rules over-redact), configuration-heavy, and text must leave the customer's trust boundary.
Source: Competitor site + tonic.ai published benchmark

Protecto

Primary Medium
protecto.ai
AI-native privacy layer built specifically for LLM pipelines, using context-preserving tokenization so masked PII and PHI does not degrade model performance — the closest positional twin to Textual's GenAI narrative and an aggressive content marketer in the same comparison queries. Materially smaller company with thinner enterprise references and no structured-data or test-data counterpart.
Source: Competitor site + search co-occurrence; tier not confirmed against deal data

Skyflow

Primary Medium
skyflow.com
Data privacy vault that tokenizes sensitive values before they reach an LLM and detokenizes authorized responses on the way back — the leading answer for runtime governance and policy-based access control in live AI workflows. Vault-centric architecture is strongest on structured PII and weakest on bulk de-identification of documents, transcripts, and free text, which is exactly Textual's center of gravity.
Source: Competitor site + search co-occurrence; tier not confirmed against deal data

Secondary Competitors — 5

Azure AI Language

Secondary Medium
azure.microsoft.com
Microsoft's managed PII detection endpoint, effectively free-at-the-margin for organizations already on an Azure commitment and the default suggestion inside Microsoft Fabric shops. Benchmarked below Textual on domain-specific text and offers no synthesis, dataset UI, or file-level review workflow — usually the incumbent to displace rather than a considered purchase.
Source: Competitor site + tonic.ai published benchmark

John Snow Labs

Secondary Medium
johnsnowlabs.com — Spark NLP for Healthcare
Healthcare-NLP specialist whose de-identification models are the entrenched choice for clinical notes, DICOM, and life-sciences text, with published accuracy validated across billions of patient notes. Deepest medical domain expertise of anyone in the set, but Spark-based and data-science-heavy to operate, and essentially irrelevant outside healthcare — it shows up in Tonic's healthcare deals and nowhere else.
Source: Competitor site and published accuracy research

Nightfall AI

Secondary Medium
nightfall.ai
Cloud-native DLP that scans SaaS apps, endpoints, and AI prompts for sensitive data and blocks or quarantines it in real time. Sits in the security team's budget solving detection-and-prevention, not the data team's problem of producing a reusable de-identified corpus — but it surfaces constantly in "PII detection for AI" searches and gets recommended alongside Textual.
Source: Category listings and review platforms

BigID

Secondary Medium
bigid.com
Enterprise data security posture and discovery platform that inventories and classifies PII across cloud, on-prem, and SaaS estates — frequently already deployed, which makes "can't BigID just do this?" a real objection. Discovery- and governance-first with batch scanning; shallow on document-level redaction and offers no synthetic replacement, so it maps the problem rather than producing AI-ready data.
Source: Category listings and review platforms

Unstructured

Secondary Medium
unstructured.io
LLM-native document ETL — 60+ file formats parsed, partitioned, chunked, and enriched into JSON ready for a vector database, with Databricks, IBM, and NVIDIA backing. Competes directly with Tonic Textual Pipelines on the extraction half of the workflow and is the better-known name for RAG preprocessing, but has no de-identification, synthesis, or compliance story at all.
Source: Category listings; may be a complement rather than an alternative

→ Validate Three questions, in order of how much they move the audit. First: in lost or stalled deals, how often is the real alternative "we'll just use Amazon Comprehend or Google DLP" versus a named privacy vendor like Limina or Protecto? Two of six primary slots currently go to hyperscaler APIs; if those are rarely serious contenders, we free roughly a dozen head-to-head queries and hand them to Limina and Protecto instead. Second: Protecto and Skyflow are tiered primary on category overlap and search co-occurrence, not on deal evidence — Protecto in particular is a heavy content marketer whose SEO footprint may exceed its real presence in your pipeline. If either rarely appears in an actual evaluation, moving them to secondary reallocates six to eight queries each. Third: is anyone here irrelevant, and is anyone missing? Unstructured is our leading removal candidate — a buyer may evaluate it as a complement to Textual Pipelines rather than an alternative to it — and any vendor that shows up in your deals but not on this page is currently invisible to the audit.

Feature Taxonomy

What Buyers Evaluate

12 buyer-level capabilities — 7 strong, 4 moderate, 1 weak — rated outside-in, which is what determines which capability queries get tested and where the audit plays offense versus defense.

Sensitive Entity Detection Accuracy Strong High

How reliably does it actually find every name, SSN, account number, and medical record ID buried in messy real-world documents — and what are its precision and recall numbers on data like mine?

Context-Preserving Synthetic Replacement Strong High

Instead of blacking everything out, swap the real names and addresses for realistic fake ones so the document still reads naturally and my model still learns something from it.

Custom Entity Types & Model Tuning Strong High

Our sensitive data includes internal identifiers no off-the-shelf model has ever seen — can I teach it to catch those without hiring an ML team?

Document, Image & File Format Coverage Strong High

My sensitive data is in scanned PDFs, Word docs, PowerPoints, spreadsheets, JSON blobs and screenshots — will it handle all of that, or just plain text?

Self-Hosted Deployment & Data Residency Strong High

Our security team will not approve shipping regulated patient or customer text to a third-party API — can I run this entirely inside our own VPC or on-prem?

Cloud, Lakehouse & Developer Platform Integrations Strong High

Does it plug into the S3 buckets, Databricks, Snowflake and Fabric we already run, and is there a real Python SDK and REST API so this can live in our pipeline instead of being a manual step?

Compliance Evidence & HIPAA Expert Determination Support Strong High

When our auditor asks whether this data is genuinely de-identified under HIPAA or GDPR, I need documentation and a statistician's sign-off — not a vendor's marketing claim.

Audio & Call Recording De-Identification Moderate Medium

Can it transcribe and strip PII out of our contact center call recordings so we can finally use them for QA and model training?

Multilingual Detection Coverage Moderate Medium

Half our support tickets are in Spanish, German and Japanese — does detection quality hold up outside English or does it just over-redact everything?

Human-in-the-Loop Review & Defensible Audit Trail Moderate Medium

A lawyer has to sign off on these redactions — I need a review screen, comments, and a timestamped record of who redacted what before anything goes out the door.

Document Extraction & Chunking for RAG Pipelines Moderate Medium

Turn our PDF and DOCX archive into clean, chunked markdown with the tables intact so we can embed it into a vector database without writing our own parser.

Real-Time Prompt & Response Redaction for Live Agents Weak Medium

I need PII stripped inline, in milliseconds, on every prompt my customer-facing agent sends to the model — not in a batch job that runs overnight.

Which Three Lead? Seven of twelve capabilities are rated Strong: Sensitive Entity Detection Accuracy, Context-Preserving Synthetic Replacement, Custom Entity Types & Model Tuning, Document, Image & File Format Coverage, Self-Hosted Deployment & Data Residency, Cloud, Lakehouse & Developer Platform Integrations, and Compliance Evidence & HIPAA Expert Determination Support. The audit tests all 12, but competitive differentiation queries will emphasize 3. Which of these best represents where Tonic.ai actually wins deals — not where the product is objectively good, but where the buyer changes their mind?

→ Validate The ratings, then the gaps. The seven Strong ratings lean heavily on your October 2025 benchmark — aggregate F1 0.96 / precision 0.97 / recall 0.96 against Comprehend, Azure AI Language, Google Cloud DLP, Presidio, GLiNER, and spaCy — which is vendor-published, and we have read it as such. Two ratings we are least sure of: Audio & Call Recording De-Identification and Multilingual Detection Coverage are rated Moderate on relative rather than absolute grounds, because your comparison page asserts Limina lacks audio entirely while Limina's own site claims audio, DICOM, and 52 languages. One of those two public claims is stale — which languages are production-grade versus best-effort, and which audio formats are GA? And the single most consequential rating in this document: Real-Time Prompt & Response Redaction is the only Weak, because the Textual page promises to "secure real-time agentic workflows" but every path we could verify — datasets, pipelines, Python SDK, REST API, MCP server — is batch or request/response, not an inline low-latency proxy with policy enforcement. If a production runtime path exists with a published latency figure, tell us and we generate runtime head-to-heads against Skyflow, Protecto, and Nightfall; if not, we treat runtime as conceded and spend those queries on batch de-identification where you win. Finally: is any capability here really two capabilities, or two really one?

Pain Points

What Buyers Are Actually Frustrated By

12 pain points — 6 high severity, 6 medium — written in the buyer's own words, because that phrasing is literally how the audit's queries get worded.

AI roadmap stuck in privacy review High High

"We have the models, the budget, and the use case — but legal won't let us touch the actual data, so our AI roadmap has been stuck in privacy review since last quarter."

AI and product teams have a funded GenAI roadmap but cannot get legal or privacy approval to use real customer documents, clinical notes, or support transcripts for fine-tuning and RAG, so projects sit in review for months.

Personas: Head of AI Platform Engineering · Director of Privacy & Regulatory Compliance · VP of Engineering · CISO

Over-redaction destroys the training corpus High High

"By the time our redaction script finishes, every document is a wall of [REDACTED] and the model can't learn anything from it. We traded a privacy problem for a garbage-data problem."

Regex and pattern-matching redaction tools black out so much of the text — including non-sensitive terms — that the resulting corpus loses the semantic context models need, making it useless for training or retrieval.

Personas: Director of Data Science & ML · Head of AI Platform Engineering

Missed PII gets baked into model weights High Medium

"If one patient name slips through into the embeddings, I can't un-ring that bell. There's no delete button on a fine-tuned model, and that's a breach notification."

A single identifier missed during de-identification ends up embedded in a vector index or baked into model weights, where it cannot be surgically removed and can be surfaced by a user prompt — turning a recall gap into a reportable breach.

Personas: CISO · Director of Privacy & Regulatory Compliance · Head of AI Platform Engineering

DIY redaction became an unfunded ML project High Medium

"We built our own PII scrubber on Presidio a year ago. Now two engineers spend every sprint patching false negatives instead of building the actual product."

Teams that started with Presidio, spaCy, or hand-rolled regex now own an unfunded ML maintenance project — retraining recognizers, adding entity types, and building evaluation harnesses — instead of shipping product.

Personas: Head of AI Platform Engineering · VP of Engineering · Director of Data Science & ML

Regulated data cannot leave the trust boundary High High

"Security shot down Comprehend and Google DLP in the first meeting. Our regulated text is not leaving our VPC, full stop — so whatever we buy has to run inside our own environment."

Managed cloud PII APIs require sending regulated text outside the organization's trust boundary, which fails security review, data residency rules, or BAA requirements — eliminating the cheapest options before evaluation even starts.

Personas: CISO · Director of Privacy & Regulatory Compliance · VP of Engineering

Manual redaction does not scale and is not safe High Medium

"Three paralegals spent six weeks redacting one production set by hand, and we still found a phone number that got through. This does not scale and it terrifies me."

Analysts and paralegals redact documents by hand for FOIA responses, litigation production, and partner data sharing — weeks of effort per batch, with any single human miss becoming an unauthorized disclosure.

Personas: Manager, Records & Information Governance · Director of Privacy & Regulatory Compliance

PII hidden inside scanned documents and images Medium Medium

"Half our archive is scanned faxes and photographed forms. Our text extractor just skips them, so we have no idea what PII is sitting in there."

Sensitive values live inside scanned PDFs, images, slide decks, and email attachments where naive text extraction either fails outright or silently drops content, leaving PII undetected in the archive.

Personas: Head of AI Platform Engineering · Manager, Records & Information Governance · Director of Data Science & ML

Call recordings locked up and unusable Medium Medium

"We have four years of call recordings that would make our support AI dramatically better, and every one of them has a credit card number read out loud in it."

Contact center recordings and transcripts are the richest source of customer insight and support-agent training data, but sit unused because they are dense with names, card numbers, and account details.

Personas: Director of Data Science & ML · Director of Privacy & Regulatory Compliance · Head of AI Platform Engineering

No defensible evidence for the auditor Medium Medium

"Our auditor asked me to prove this data is de-identified under HIPAA. I had a vendor datasheet and nothing else. That is not an answer I want to give twice."

Organizations cannot demonstrate to an auditor or regulator what was detected, what was transformed, who reviewed it, or what residual re-identification risk remains — so de-identification claims rest on vendor assertion rather than evidence.

Personas: Director of Privacy & Regulatory Compliance · CISO · Manager, Records & Information Governance

Structured and unstructured tool sprawl Medium Medium

"We have one tool masking the database and a different one scrubbing the PDFs, and they replace the same patient's name with two different fake names. Nothing reconciles."

Separate vendors handle database test data and document or free-text de-identification, producing two contracts, two security reviews, and inconsistent treatment of the same individual across structured and unstructured sources.

Personas: VP of Engineering · CISO · Head of AI Platform Engineering

Batch redaction does not protect live agents Medium Medium

"Our training data is squeaky clean, but every customer chat goes straight to the model with their real account number in it. The batch job doesn't help me at runtime."

Batch de-identification cannot sit in the request path of a customer-facing agent, so live prompts and tool outputs reach the model unprotected while the offline corpus is carefully sanitized.

Personas: Head of AI Platform Engineering · CISO · VP of Engineering

Non-English data excluded from AI projects Medium Medium

"Our EU and APAC data just gets excluded from every AI project because nobody trusts the redaction on German or Japanese text. That's a third of our customers we can't learn from."

Detection quality degrades outside English, so global organizations either exclude non-English data from AI initiatives entirely or accept aggressive over-redaction that renders it unusable.

Personas: Director of Privacy & Regulatory Compliance · Director of Data Science & ML · Head of AI Platform Engineering

→ Validate Severity, phrasing, and one deliberate omission. Six of twelve pains are rated High, which means half this set competes for the same "test first" slot — if you had to name the three that actually stall deals, are they the legal-approval blocker, the Presidio maintenance burden, and the trust-boundary rule, or is that ordering wrong? On phrasing: the buyer language is what the audit's prompts get written in, so if "we traded a privacy problem for a garbage-data problem" is not how your buyers talk, tell us the sentence they actually say. And on omissions — we deliberately left out a pricing or TCO pain even though G2 reviews cite cost and setup complexity, because those complaints attach to Structural's database-masking workflow rather than Textual, which sells pay-as-you-go per 1,000 words. If Textual-specific pricing objections are real in your deals, that becomes a thirteenth pain point and a whole query cluster. Two others worth considering: quantifying residual re-identification risk (if buyers ask for a statistical risk number, not just a process), and throughput on very large archives (if "how long to process ten million documents" is a live objection rather than an afterthought).

Layer 1 Technical Findings

What the Crawl Found

50 pages of tonic.ai analyzed at the HTML level — crawler directives, sitemap, schema markup, heading structure, and date signals. 15 findings: 13 diagnostic, 2 requiring manual verification. Zero critical.

Engineering — Actionable Now There is no critical blocker here: tonic.ai serves complete server-rendered HTML on all 50 pages, robots.txt permits every declared AI crawler, and schema coverage sits at 0.84. The four high-severity items are all engineering-owned and all shippable in under a day each. Start with the Cloudflare check — every response carries server: cloudflare, and Cloudflare's Bot Fight Mode and "Block AI bots" managed rules operate at the edge, above robots.txt. Cloudflare changed its default to block AI crawlers in July 2025 and its network handles roughly 24% of all web traffic (Cloudflare, July 2025), so a permissive robots.txt is not proof of access; confirm from logs that GPTBot, ClaudeBot, PerplexityBot, and Google-Extended got 200s over the last 30 days. If they did not, that supersedes everything else in this document. Then: remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template, which currently tells a reader that the Limina and K2View pages are 2023 comparisons of Delphix; and emit <lastmod> from the Webflow CMS across all 1,878 sitemap URLs, which today publish none. The two Content-owned findings below are noted for planning, not for this sprint.

🟡 Competitor comparison pages carry stale, contradictory, and wrong-competitor "Last updated" stamps

What we found: All three /alternatives/ pages emit a JSON-LD dateModified of 2026-07-16, but the visible body text says something different on each: /alternatives/limina reads "Last updated August 2025", /alternatives/k2view reads "Last updated December 2025", /alternatives/delphix reads "Last updated July 2026". Worse, all three also render a second, orphaned block reading "Book a DemoLast updated December 2023. Comparison based on Delphix's full suite of services." — so the Limina page and the K2View page both tell a reader that they are a December 2023 comparison of Delphix. The /vs/ pages are clean (single "Last updated July 2026" stamp).

Why it matters: Comparison pages are the single most-cited page type in vendor-evaluation queries, and an LLM reads the visible text, not the JSON-LD. On the Limina page — the head-to-head against the closest competitor in the knowledge graph — the strongest recency signal a model can extract is "December 2023", and the strongest topical signal in that same sentence names the wrong competitor. That is both a freshness penalty and an accuracy defect on the highest-value asset in the competitive set.

Business consequence: When a buyer asks an assistant "Tonic Textual vs Private AI" or "Limina alternatives for PHI redaction," the page built to win that answer identifies itself as a two-and-a-half-year-old comparison of a different vendor — so the answer gets assembled from whoever describes the matchup coherently instead.

Recommended fix: Remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template (it is clearly an un-cleared static element in the component). Then bind the visible "Last updated" string to the same CMS field that feeds JSON-LD dateModified so the two can never diverge, and re-stamp Limina and K2View to their true review dates.

Impact: high Effort: < 1 day Owner: Engineering Affected: /alternatives/limina, /alternatives/delphix, /alternatives/k2view

🟡 Comparison content is published at two competing URLs with conflicting canonical handling

What we found: The Synthesized comparison exists at both /vs/synthesized (2,460 words) and /blog/tonic-vs-synthesized (2,450 words) with the same H1 and near-identical body, and each page self-canonicalizes. Separately, /alternatives/delphix sets rel=canonical to /vs/delphix, yet both URLs are submitted in sitemap.xml as independent entries, and a third Delphix comparison is reachable from /vs/delphix's own "Full comparison article" link. The Delphix comparison therefore occupies three URLs and the Synthesized comparison two.

Why it matters: Splitting one comparison across two self-canonical URLs divides link equity and retrieval signals between them, and gives retrieval-augmented systems two competing candidates for the same query with no signal about which is authoritative. Submitting a URL in the sitemap that canonicalizes elsewhere sends contradictory instructions to crawlers and wastes crawl budget on a page the site itself says is not the original.

Business consequence: Two self-canonical versions of the same comparison split the retrieval signal for "Tonic vs Synthesized" queries, so neither URL accumulates the authority a single canonical page would — on exactly the page type that decides late-stage de-identification evaluations.

Recommended fix: Pick one canonical home per competitor — the /vs/{competitor} path is the stronger pattern — and 301 the duplicates to it, or at minimum set rel=canonical on /blog/tonic-vs-synthesized to /vs/synthesized. Remove any URL that canonicalizes elsewhere (currently /alternatives/delphix) from sitemap.xml.

Impact: high Effort: < 1 day Owner: Engineering Affected: /vs/synthesized, /blog/tonic-vs-synthesized, /vs/delphix, /alternatives/delphix

🟡 sitemap.xml provides no <lastmod> on any of its 1,878 URLs

What we found: https://www.tonic.ai/sitemap.xml is a flat (non-index) sitemap listing 1,878 URLs. Every entry contains only a <loc> element — zero URLs carry <lastmod>, and zero carry <priority> or <changefreq>. The HTTP Last-Modified header is no substitute: the origin returned last-modified: Mon, 27 Jul 2026 16:36:02 GMT — the moment of our request — for pages whose content is years old, so it is a build timestamp, not a content timestamp.

Why it matters: With no <lastmod> and a header that always reports "now", a crawler has no machine-readable way to tell which of 1,878 pages changed. Recrawl scheduling falls back to heuristics, so genuinely refreshed pages — including the comparison pages Tonic actively maintains — wait in the same queue as a 2021 careers post. Freshness is a documented driver of AI citation selection, and the site currently forfeits the cheapest signal available to it.

Business consequence: The benchmark and comparison pages you actively maintain — the assets that answer "most accurate PII detection tool" and "best de-identification software for LLM training" — queue for recrawl behind 1,443 changelog entries, so a refresh reaches AI indexes late or not at all.

Recommended fix: Emit <lastmod> from the Webflow CMS "last published" field for every URL in sitemap.xml, and stop returning a request-time Last-Modified header on cached HTML (serve the origin's real content timestamp, or omit the header and rely on ETag).

Impact: high Effort: 1-3 days Owner: Engineering Affected: sitemap.xml — all 1,878 URLs sitewide

🔵 Release notes are 77% of the sitemap, burying the 435 commercially relevant URLs

What we found: Of the 1,878 URLs in sitemap.xml, 1,443 (76.8%) are per-release changelog entries: 820 under /tonic-structural-release-notes, 447 under /tonic-textual-release-notes, 172 under /tonic-fabricate-release-notes, and 4 under /product-release-notes. Only 435 URLs are commercial or editorial content. The sitemap is a single flat file with no index and no segmentation.

Why it matters: Release-note pages are near-duplicates of one another, carry almost no standalone text, and are essentially never the right answer to a buyer question — yet they outnumber the product, comparison, capability, and case-study pages by more than three to one in the only prioritization signal the site publishes. Crawl budget spent re-fetching 1,443 changelog entries is crawl budget not spent on the benchmark and comparison pages that actually win citations.

Business consequence: Crawl budget consumed by per-release changelog entries is budget not spent on the 435 pages that answer questions like "de-identify clinical notes for LLM training," which slows how quickly new commercial content becomes citable at all.

Recommended fix: Convert sitemap.xml into a sitemap index with separate child sitemaps (pages, blog, guides, case studies, release notes) so crawlers can prioritize, and consider consolidating per-release entries into one paginated changelog per product rather than one indexable URL per release.

Impact: medium Effort: 1-3 days Owner: Engineering Affected: sitemap.xml; the 1,443 release-note URLs

🔵 BlogPosting and case-study schema omit dateModified, so refreshed posts cannot signal recency

What we found: Every /blog/ post and every /case-study/ page we inspected emits datePublished in its BlogPosting/Article JSON-LD but no dateModified — 11 of 11 blog and case-study pages in the inventory. The /vs/, /alternatives/, /ai-model-benchmarks/ and guide templates do emit both. The result is that the only machine-readable date on a blog post is the day it first shipped, and pages like /blog/redacting-sensitive-free-text-data-build-vs-buy (published 2023-12-12) and /blog/how-to-de-identify-legal-documents-with-tonic-textual (2024-04-15) present as 958 and 833 days old respectively even if they have been edited since.

Why it matters: Adding dateModified is the difference between a refresh being invisible and a refresh being credited. Without it, the content team can rewrite a post and gain nothing in freshness-weighted retrieval, which removes the incentive to maintain the archive at all. It also makes the stale-content problem below impossible to fix without republishing under a new date.

Business consequence: A refreshed build-vs-buy post still presents as 2023 to any system weighing recency, so the content team's maintenance work on de-identification topics earns no credit in exactly the queries where recency breaks a tie between comparable vendors.

Recommended fix: Add a dateModified property to the BlogPosting/Article JSON-LD in the blog and case-study templates, bound to the CMS "last updated" field, and surface the same value as a visible "Updated {date}" line beneath the byline.

Impact: medium Effort: 1-3 days Owner: Engineering Affected: Blog post template (149 URLs) and case-study template (18 URLs)

🔵 Three commercially important pages are more than a year old and two more are past six months

What we found: Dated content-marketing pages in the inventory that fall outside the citation window: /blog/redacting-sensitive-free-text-data-build-vs-buy (2023-12-12, 958 days), /blog/how-to-de-identify-legal-documents-with-tonic-textual (2024-04-15, 833 days), /case-study/tax-agency-unblocked-ai-llm-workflows-with-secure-data-synthesis (2025-05-02, 451 days), /case-study/ontra-accelerates-ai-innovation-in-legal-tech-with-tonic-textual (2025-11-11, 258 days), and /blog/self-serve-custom-entity-types-in-tonic-textual (2025-11-17, 252 days). Separately, the three /playbooks/ pages publish no date at all in either the rendered body or JSON-LD. The two >365-day posts are substantive (content depth 0.80 and 0.75) and are the site's primary build-vs-buy and document-redaction walkthroughs.

Why it matters: These are not low-value pages being allowed to age — the build-vs-buy post is the natural landing point for the "we built our own PII scrubber on Presidio" buyer, and the legal-documents walkthrough is the only step-by-step de-identification tutorial on the site. Both now describe a product that has since gained Custom Entity Types, audio, an MCP server, and LLM-assisted NER, so they are simultaneously the stalest and the least accurate pages in the set.

Business consequence: The build-vs-buy post is where the "two engineers patch Presidio false negatives every sprint" buyer lands, and it currently answers that query by describing a version of Textual that predates Custom Entity Types, audio, and the MCP server.

Recommended fix: Refresh the two >365-day posts against current Textual capability and republish with dateModified set (see the schema finding above). Add a visible publish/updated date to the three /playbooks/ pages, which currently carry no date signal of any kind.

Impact: medium Effort: 1-2 weeks Owner: Content Affected: 5 dated pages plus 3 undated /playbooks/ pages

🔵 Heading structure breaks down on several high-traffic commercial pages

What we found: Every page in the inventory has exactly one H1, which is good. The problems are below it. /integrations nests 96 H3 headings under just 2 H2s, so the connector directory has effectively no hierarchy. The three /playbooks/ pages use 6 H2s and zero H3s, leaving a 35-entry entity-type glossary — roughly half the page's word count — completely unheaded. /products/textual uses a full two-sentence marketing paragraph ("Instantly redact sensitive PII from text and audio to safely train models, and secure real-time agentic workflows. Get the essential privacy layer for every LLM interaction.") as an H2. /pricing repeats the generic headings "Features", "Common questions", and "Enterprise" three times each, once per product block, with nothing distinguishing them. The homepage carries 94 H3s, nearly all of which are resource-card titles rather than page structure.

Why it matters: Headings are how a retrieval system segments a page into citable passages and how it labels what each passage is about. A 96-item flat list, an unheaded glossary, a sentence used as a heading, and three identical "Features" headings all produce passages that either cannot be delimited or cannot be told apart, which pushes otherwise useful content below the citation threshold.

Business consequence: A 96-connector flat list and three identical "Features" headings give a retrieval system no way to isolate the passage that answers "does Tonic Textual work with Snowflake" or "Tonic Textual pricing," so those passages go unquoted even though the content exists.

Recommended fix: Group /integrations connectors under category H2s (Relational, NoSQL, Warehouse/Lakehouse, Files, SaaS) with connector names as H3. Give the playbook entity glossary an H2 and group entities by class. Replace the /products/textual paragraph-H2 with a short noun phrase and move the sentence into body copy. Qualify the repeated /pricing headings ("Fabricate features", "Structural common questions", and so on).

Impact: medium Effort: 1-3 days Owner: Engineering Affected: /integrations, 3 /playbooks/ pages, /products/textual, /pricing, homepage

🔵 Nine commercial pages ship only the site-wide Organization block with no page-specific schema

What we found: Schema coverage across the site is generally strong — 33 of 50 inventoried pages carry an appropriate specific type (Product, SoftwareApplication, FAQPage, TechArticle, Review, HowTo, CollectionPage). Nine do not: all four /partners/ pages, all three /playbooks/ pages, /integrations, and /products/validate emit only the global Organization / WebSite / Person / PostalAddress / ContactPoint block that appears in the site footer template, with no WebPage, Product, HowTo, or BreadcrumbList of their own. Separately, docs.tonic.ai/textual emits no JSON-LD at all and has no meta description.

Why it matters: Those nine pages carry the marketplace, lakehouse, and hyperscaler integration story — exactly the material a buyer query like "PII redaction inside Snowflake" should retrieve — and the playbooks are the only procedural how-to content on the marketing site. Without HowTo, Product, or even BreadcrumbList markup, they are unlabeled prose in a corpus where their siblings are richly typed, and they lose the breadcrumb trail that helps a model place them in the site hierarchy.

Business consequence: Queries like "PII redaction inside Snowflake" or "how to redact audio recordings" should land on the partner and playbook pages, and those nine pages arrive at the retrieval layer untyped while the rest of the site arrives labeled.

Recommended fix: Add BreadcrumbList and WebPage to the /partners/, /playbooks/, /integrations and /products/validate templates; add HowTo to the playbook template (each already has a numbered "Playbook steps" section that maps directly onto HowToStep); add SoftwareApplication to /products/validate. On docs.tonic.ai, add a meta description and TechArticle markup to the GitBook theme.

Impact: medium Effort: 1-3 days Owner: Engineering Affected: 4 /partners/ pages, 3 /playbooks/ pages, /integrations, /products/validate, docs.tonic.ai

🔵 Eight commercially relevant pages are too thin to sustain a citation

What we found: Body word counts after stripping nav and footer: /products/validate 255 words (content depth 0.30), /partners/aws 319, /partners/databricks 325, /partners/microsoft/fabric 383, /products/tonic-datasets 444, /partners/snowflake/textual-native-app 459 (all 0.40–0.50), plus /integrations at 1,684 words that are almost entirely one-sentence connector blurbs (0.45) and /playbooks/audio-redaction-and-synthesis where roughly 380 of 623 words are a repeated entity-type glossary shared verbatim with two other playbooks (0.40). For comparison, the same site's /ai-model-benchmarks/textual-benchmark page runs 6,424 words with per-competitor precision/recall/F1 tables.

Why it matters: Every one of these thin pages sits on a query a buyer actually asks — "Tonic Textual on Databricks", "redact PII in Snowflake", "de-identify audio recordings". A 320-word page of two-sentence claims gives a retrieval system nothing specific enough to quote, so the query is answered from a competitor's or a hyperscaler's documentation instead. The gap is especially stark because Tonic clearly knows how to write the deep version — it did so for the benchmarks.

Business consequence: A 319-word partner page cannot supply a quotable specific for "Tonic Textual on Databricks," so that de-identification query gets answered from hyperscaler documentation rather than from tonic.ai.

Recommended fix: Bring each partner page up to a concrete implementation narrative: the specific integration mechanism, a worked example, throughput or accuracy figures where they exist, and the setup steps. Replace the shared entity glossary on the playbooks with a single canonical entity-types reference page and link to it. Expand /integrations connector entries with supported versions, file types, and limits.

Impact: medium Effort: 2-4 weeks Owner: Content Affected: /products/validate, /products/tonic-datasets, 4 /partners/ pages, /integrations, /playbooks/audio-redaction-and-synthesis

🔵 "No items found." placeholders render on five live pages, including two competitor comparisons

What we found: The Resources module renders the literal string "No items found." in the served HTML of /alternatives/delphix, /alternatives/k2view, /capabilities/expert-determination, /capabilities/government-redaction, and /case-study/ontra-accelerates-ai-innovation-in-legal-tech-with-tonic-textual. On /alternatives/delphix the module heading promises "Further reading and resources on Delphix alternatives" and is followed immediately by the empty-state string.

Why it matters: This is text a crawler ingests as page content. It adds an empty-state string to pages that are otherwise arguing a competitive case, and it removes the internal links that would normally pass authority from a comparison page into the supporting benchmark and blog content. On the expert-determination and government-redaction pages it also strips the only outbound paths to the deeper compliance material.

Business consequence: "No items found." is ingested as body text on two live competitor comparisons and on the expert-determination page that supports HIPAA de-identification queries, while the internal links that would carry authority into the benchmark content are simply absent.

Recommended fix: Populate the Resources collection on these five pages with existing relevant content — the Textual benchmark and PrivacyBench pages for the comparison pages, the HIPAA expert-determination blog posts for /capabilities/expert-determination, the FOIA redaction post for /capabilities/government-redaction — and change the Webflow component to hide the whole module when the collection is empty rather than printing the placeholder.

Impact: medium Effort: < 1 day Owner: Engineering Affected: 5 pages across /alternatives/, /capabilities/, and /case-study/

⚪ Repeated sentences and a broken statistic render on two solution and partner pages

What we found: /partners/microsoft/fabric renders "Choose an output OneLake folder for storing the de-identified files." three times consecutively and "If needed, adjust settings and reprocess." four times consecutively inside its numbered setup sequence — on a page that only has 383 words of body copy, roughly 10% of the page is a duplicated sentence. /solutions/use-case/llm-privacy-proxy displays its headline metric as "<0 min — Real-time redaction", which reads as "less than zero minutes". /ai-training-data renders its "On this page" block and "Generate the AI training data you need" CTA twice in a row.

Why it matters: These are small but they land on the two pages most relevant to the knowledge graph's weakest feature area. "<0 min" is the only latency claim anywhere on the site for real-time redaction, and as rendered it is not a number a model can quote or a buyer can trust. The Fabric repetition makes the shortest integration page look automated rather than authored.

Business consequence: "<0 min" is the only latency figure published anywhere on the site — the exact number a buyer asking "real-time PII redaction latency" needs to compare you against a runtime vendor — and as rendered it is unquotable.

Recommended fix: Remove the duplicated CMS repeater rows on /partners/microsoft/fabric and /ai-training-data. Replace the "<0 min" stat on /solutions/use-case/llm-privacy-proxy with a real, defensible latency figure and its measurement conditions, or remove the stat block.

Impact: low Effort: < 1 day Owner: Content Affected: /partners/microsoft/fabric, /solutions/use-case/llm-privacy-proxy, /ai-training-data

⚪ Seven pages have no og:image, including two competitor comparison pages

What we found: og:title and og:description are present sitewide, but og:image is absent on /faqs, /vs/delphix, /vs/synthesized, /partners/aws, /partners/databricks, /partners/microsoft/fabric, and /partners/snowflake/textual-native-app. Meta descriptions are present on all 49 www pages inspected (93–254 characters) and missing only on docs.tonic.ai/textual.

Why it matters: Low impact on retrieval itself, but /vs/delphix and /vs/synthesized are the pages most likely to be shared into Slack, LinkedIn, and sales email during an active evaluation, and they currently unfurl without a preview image. It is a cheap fix on the pages where social amplification matters most.

Business consequence: The two comparison pages your sales team pastes into Slack and email during a live de-identification evaluation are the two that unfurl as a bare link.

Recommended fix: Add an og:image to the /vs/, /partners/, and /faqs templates, defaulting to the existing Tonic.ai social card where no bespoke image exists. Add a meta description to docs.tonic.ai/textual.

Impact: low Effort: < 1 day Owner: Marketing Affected: /faqs, /vs/delphix, /vs/synthesized, 4 /partners/ pages, docs.tonic.ai/textual

⚪ A markdown artifact has produced one malformed href on the homepage

What we found: The homepage contains href="https://textual.tonic.ai/signup](https://textual.tonic.ai/signup" — a markdown link fragment ("](") pasted into a Webflow link field, producing a single URL with the closing-bracket sequence embedded. This is the only malformed href we found across the 50 pages fetched, and a 40-URL random sample of non-release-note sitemap entries returned HTTP 200 for all 40, so there is no broader broken-link problem.

Why it matters: Minor on its own — one dead signup path on the highest-traffic page — but worth fixing because it is a conversion link and because it indicates content is being pasted from markdown into the CMS without validation, which is the same class of error that produced the orphaned "Last updated December 2023" block on the comparison pages.

Business consequence: The homepage signup path is the shortest route from an AI-referred visitor to a Textual trial, and it currently dead-ends.

Recommended fix: Correct the link to https://textual.tonic.ai/signup, and add a build-time link validation step (or a Webflow pre-publish check) that rejects hrefs containing "](" or whitespace.

Impact: low Effort: < 1 day Owner: Engineering Affected: Homepage (https://www.tonic.ai/)

Manual Verification Checklist

The following items could not be assessed through our analysis method (rendered markdown). We recommend your engineering team verify these manually before the validation call.

robots.txt allows every AI crawler — confirm Cloudflare is not blocking them above it High

What to check: https://www.tonic.ai/robots.txt contains a single User-agent: * group that disallows only /styleguide, so GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Googlebot and Bytespider are all permitted. docs.tonic.ai/robots.txt goes further and publishes Content-Signal: ai-train=yes, search=yes, ai-input=yes with Allow: /. However, every response from www.tonic.ai carries server: cloudflare, and Cloudflare's Bot Fight Mode, AI Labyrinth, and "Block AI Bots" managed rules operate at the edge, independent of robots.txt — a site can publish a fully permissive robots.txt and still return 403s to GPTBot and ClaudeBot. We fetched with a browser user-agent and cannot determine from the outside how the edge treats declared AI user-agents. This is the single highest-leverage technical check on the site: if the WAF is silently blocking AI crawlers, nothing downstream can produce citations, and confirming 200s converts this from a risk into a verified strength worth stating in the audit.

Recommended action: In the Cloudflare dashboard, check Security > Bots for Bot Fight Mode, the "Block AI bots" managed rule, and any custom WAF/rate-limit rules matching AI user-agents; then confirm against origin/Cloudflare logs that GPTBot, ClaudeBot, PerplexityBot and Google-Extended requests returned 200 over the last 30 days. Explicitly allowlist the seven crawlers if any rule intercepts them.

Effort: < 1 day Owner: Engineering Affected: All of www.tonic.ai (1,878 sitemap URLs) plus docs.tonic.ai

Confirm the JavaScript-rendered DOM matches the server HTML we analyzed Low severity — confirmation step

What to check: This analysis parsed the raw server-rendered HTML returned by the origin, which let us assess schema markup, meta tags, OG tags, and content directly rather than inferring them. All 50 pages returned full body content, headings, and JSON-LD without executing JavaScript, so no page shows client-side-rendering symptoms. What we could not check is the opposite direction: whether the browser-rendered DOM adds or reorders content the server HTML does not contain. The homepage's four-panel tabbed hero, the /integrations and /customers filter widgets, and the FAQ accordions are interactive components whose full text we did find in the server HTML, but whose behaviour under a JS-executing crawler we did not observe. Server-rendered HTML is the stronger position to be in and this is a confirmation step, not a suspected defect — but a mismatch between server HTML and rendered DOM is the one failure mode our method structurally cannot see.

Recommended action: Spot-check five pages (homepage, /products/textual, /integrations, /pricing, /vs/synthesized) in Google Search Console's URL Inspection "View crawled page" and in a headless-Chrome crawler, comparing rendered text and JSON-LD against view-source. Confirm the homepage hero's non-active tab panels and the /integrations connector list are present without JavaScript.

Effort: < 1 day Owner: Engineering Affected: Interactive components on the homepage, /integrations, /customers, and FAQ accordions sitewide

Site Analysis Summary

Pages analyzed 50 of 435 commercially relevant sitemap URLs
Commercially relevant pages in sample 50 of 50
Avg heading hierarchy 0.765
Avg content depth 0.632
Avg passage extractability 0.692
Avg schema coverage 0.84
Freshness (weighted) 0.606 — content marketing 0.606 (24 pages); product/commercial unable to assess (23 unscored); structural/reference unable to assess (3 unscored)
Freshness page tiers (content marketing) 12 under 90d · 8 over 180d · 3 over 365d
Findings by severity 0 critical · 4 high · 7 medium · 4 low
Crawler directives 7 of 7 declared AI crawlers allowed in robots.txt

Partial Sample This analysis covers 50 pages of the 435 commercially relevant URLs in sitemap.xml — roughly 12%, weighted toward product, comparison, capability, partner, and case-study pages. Sitewide claims about schema, headings, and thin content are extrapolated from that sample and should be read as representative, not exhaustive. Freshness is a bigger caveat: 26 of the 50 pages could not be scored at all because they publish no date in either the rendered body or JSON-LD — 23 product and commercial pages and all 3 structural/reference pages — so the 0.606 weighted score reflects only the 24 dated content-marketing pages. Fixing the <lastmod> and dateModified findings above would make the next measurement materially more complete.

Next Steps

What Happens Now

Why Now

  • Buyer discovery is shifting quarter over quarter — 94% of B2B buyers now use LLMs somewhere in the buying process (6sense, November 2025), and crawler demand is scaling with it: Akamai measured 1.6 billion daily AI bot requests across its CDN, up 78% over six months (Akamai, February 2026).
  • Presence compounds. Brand web mentions predict AI citation far more strongly than backlinks or organic traffic — roughly three times as predictive as backlinks (Seer Interactive, October 2025) — so the visibility you build now feeds the signal that drives later citations.
  • Visibility in a narrow category concentrates: top brands in narrow categories appeared in 70–90% of repeated AI recommendation runs, even though the exact recommendation list is almost never identical twice (SparkToro & Gumshoe, January 2026, consumer categories).
  • Unstructured de-identification is still early-innings in GEO. Right now the competition for "how do I de-identify documents for LLM training" is mostly inaction — not entrenched strategy. That is a narrower window than it looks.

Once the inputs on these pages are confirmed, the audit measures how Tonic.ai actually appears across the selected AI platforms for the questions your buyers ask — "how do I de-identify clinical notes for LLM training," "best alternative to Private AI for PHI redaction," "build vs buy a PII scrubber on Presidio," "can I run PII redaction inside my own VPC." You will see exactly which of those queries return answers naming Limina, Presidio, Comprehend, or Skyflow but not Tonic.ai, which name you but describe you wrongly, and what specifically it would take to appear in the ones you are missing. Fixing the Layer 1 items in the meantime means the audit measures a repaired baseline rather than a known-broken one — the Cloudflare check in particular determines whether any of the rest of it is measurable at all.

01

Validation Call

45–60 minutes. We walk through this document together, resolve the questions in the Pre-Call Checklist, and lock the competitor tiers, personas, and capability ratings the query set will be built from.

02

Query Generation & Execution

We generate buyer queries from the validated personas, pain points, and capabilities, then run them across the AI platforms selected for this engagement and capture every response and citation.

03

Full Audit Delivery

Visibility analysis, competitive positioning against the confirmed set, and a three-layer action plan — technical, content, and authority — prioritized by which gaps actually cost you citations.

Engineering Can Start Today Three Layer 1 fixes need no decision from anyone at the validation call. One: verify Cloudflare is not intercepting AI crawlers above robots.txt — check Security > Bots for Bot Fight Mode and the "Block AI bots" managed rule, then confirm from logs that GPTBot, ClaudeBot, PerplexityBot, and Google-Extended received 200s over the last 30 days. Crawler access is currently unverified, and if the edge is blocking them it supersedes everything else in this document. Two: remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template and bind the visible "Last updated" string to the same CMS field that feeds JSON-LD dateModified — a sub-day fix on the Limina page, which is the head-to-head against your closest competitor. Three: emit <lastmod> from the Webflow CMS across all 1,878 sitemap URLs and stop returning a request-time Last-Modified header. These don't depend on the rest of the audit and will improve your baseline visibility before we even measure it.

Before the Call

Your Pre-Call Checklist

Two jobs before we meet. The questions on the left require your judgment — no one knows your business better than you. The engineering tasks on the right don't require the call at all.

Questions for You
Is this audit scoped to Tonic Textual, or to the full Tonic.ai platform including Structural and Fabricate?
If wrong: Delphix, K2View, and Synthesized re-enter the competitive set and the personas, capabilities, and query vocabulary all change with them.
Is Real-Time Prompt & Response Redaction genuinely weak, or is there a production inline path with published latency?
If wrong: we generate runtime head-to-heads against Skyflow, Protecto, and Nightfall instead of conceding the runtime category and spending those queries on batch.
In lost or stalled deals, is the real alternative "we'll just use Comprehend or Google DLP," or a named vendor like Limina or Protecto?
If wrong: two of six primary slots are spent on hyperscaler APIs that never seriously compete, and roughly a dozen head-to-head queries are misallocated.
Do Protecto and Skyflow actually appear in your evaluations, and is Unstructured a competitor or a complement?
If wrong: six to eight head-to-head queries per vendor are aimed at companies you never meet in a deal.
Is Terrence Whitlock — records and legal redaction — a funded buying motion, or is Guided Redaction still early-access?
If wrong: an entire query vocabulary tests a market you aren't selling into, and should be reallocated to the AI platform and security personas.
Does VP Engineering (Sofia Halvorsen) really own the platform budget, or does it sit with the CISO or the AI platform org?
If wrong: one of our three veto-holders is wrong and the approval-stage queries are written for the wrong reader.
Which three of the seven Strong capabilities actually change a buyer's mind, rather than merely being good?
If wrong: the differentiation queries emphasize capabilities that don't decide deals, and the ones that do go under-tested.
Which audio formats are GA and which languages are production-grade — your page says Limina lacks audio, theirs claims audio, DICOM, and 52 languages?
If wrong: two Moderate ratings are mis-set, and multilingual and audio queries play offense when they should play defense (or vice versa).
Which three of the six high-severity pain points actually stall deals, and is a Textual-specific pricing objection missing from the set?
If wrong: the first wave of queries tests frustrations that don't block purchases, and a live objection goes entirely unmeasured.
Does Priya Raghunathan (Head of AI Platform Engineering) control spend, and is Ethan Brandt the entry point or a late-stage reviewer?
If wrong: Priya needs procurement-stage queries she currently doesn't get, and the over-redaction query cluster is ordered incorrectly.
Is HIPAA expert determination or GDPR residency the more common regulatory blocker, and is a clinical, DPO, or vendor-risk role missing from the persona set?
If wrong: Dana Okonkwo's cluster is written in the wrong regulatory vocabulary and an entire buying conversation goes unmeasured.
For Engineering — Start Now
Audit Cloudflare bot rules and confirm 30 days of 200s for GPTBot, ClaudeBot, PerplexityBot, and Google-Extended
robots.txt is permissive, but edge rules operate above it — this determines whether anything else in the audit is measurable.
Delete the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ template
It currently tells readers the Limina and K2View pages are 2023 comparisons of Delphix. Bind the visible date to the CMS field feeding dateModified.
Emit <lastmod> from Webflow across all 1,878 sitemap URLs and stop serving a request-time Last-Modified header
Zero URLs carry a content timestamp today, so crawlers cannot tell which pages changed.
Collapse the duplicate comparison URLs onto /vs/{competitor} and drop /alternatives/delphix from sitemap.xml
Synthesized occupies two self-canonical URLs and Delphix three; the sitemap currently submits a page that canonicalizes elsewhere.
Add dateModified to the BlogPosting/Article JSON-LD in the blog and case-study templates
167 pages currently publish only a first-ship date, so no refresh can ever be credited.
Suppress the "No items found." Resources placeholder and fix the malformed homepage signup href
Both are Webflow template hygiene: five live pages print an empty-state string, and one homepage link contains a pasted markdown fragment.
Spot-check five pages for render parity in GSC URL Inspection and a headless-Chrome crawler
Server HTML looks complete on all 50 pages; this confirms the rendered DOM doesn't diverge — the one failure mode our method can't see.
Alignment

We're Aligned On

This isn't a contract — it's a shared understanding. The audit runs against what's below. If something changes between now and the call, we adjust. The goal is to make sure we're asking the right questions for the right buyers against the right competitors.
Already Confirmed
Competitive set — 11 vendors identified: 6 primary (Limina, Microsoft Presidio, Amazon Comprehend, Google Cloud Sensitive Data Protection, Protecto, Skyflow) + 5 secondary (Azure AI Language, John Snow Labs, Nightfall AI, BigID, Unstructured)
Persona set — 6 personas: 3 decision-makers (CISO, Privacy & Compliance Director, VP Engineering), 1 evaluator (Head of AI Platform Engineering), 2 influencers (Director of Data Science, Records & Information Governance Manager)
Feature taxonomy — 12 buyer-level capabilities with outside-in strength ratings: 7 strong, 4 moderate, 1 weak
Pain point set — 12 buyer frustrations with severity ratings: 6 high, 6 medium
Layer 1 technical audit — 15 findings across 50 pages (0 critical, 4 high, 7 medium, 4 low), engineering notified via the checklist above
Entity normalization — Private AI and Limina collapse to one entity so head-to-head share against the closest competitor isn't split across both names
Decided at the Call
Audit scope — Tonic Textual standalone vs. the full Tonic.ai platform. This determines the entire competitive set; nothing else can be finalized until it is.
Real-Time Prompt & Response Redaction — is the Weak rating correct, or does a production runtime path exist? The single most consequential rating in the graph.
Competitor tiers — Protecto and Skyflow (tiered on search co-occurrence, not deal data), whether hyperscaler APIs deserve two primary slots, and whether Unstructured belongs at all
Persona corrections — is Terrence Whitlock's records/legal motion funded, and does VP Engineering or the CISO hold the budget veto?
Feature overweighting — which 3 of the 7 Strong capabilities lead the differentiation queries, plus confirmation of the Audio and Multilingual Moderate ratings
Pain point prioritization — top 3 of the 6 high-severity pains to test first, and whether a Textual-specific pricing objection joins the set
Client
Date