AI assistants are becoming the first place teams ask how to strip PII, PHI, and PCI out of documents and transcripts before that data reaches a model — and the vendors named in those answers today are locking in an advantage while the category is still being defined. Before we run the audit, we need to make sure we're asking the right questions about the right competitors to the right buyers. This document presents what we've learned about Tonic.ai's market — your job is to tell us what we got right, what we got wrong, and what we missed.
Before we measure how often Tonic.ai is cited in unstructured de-identification and redaction questions, these three signals tell us whether AI crawlers can reach the site, read it, and tell what is current. All three are derived mechanically from the Layer 1 crawl.
Unstructured data de-identification is a category being defined right now, and a large share of the buyers defining it start by asking an assistant rather than a search engine — 94% of B2B buyers now use LLMs somewhere in the buying process (6sense, November 2025). That makes the answer to "how do I de-identify documents for LLM training" a contested asset, and the vendors named in it today accumulate an advantage that compounds, because brand web mentions predict AI citation far more strongly than backlinks or organic traffic do (Seer Interactive, October 2025) — visibility feeds the signal that produces more visibility. Tonic.ai enters this with unusual raw material for the category — a published, reproducible accuracy benchmark and a self-hosted deployment story — but raw material only converts into citations if AI systems can reach it, parse it, and tell that it is current.
This Foundation Review is the input layer for that measurement, and it asks you to confirm or correct four things. The competitive landscape determines which vendors we test Tonic.ai against head-to-head and which we merely watch. The buyer personas determine the intent, vocabulary, and altitude of every query we generate. The feature and pain-point taxonomies supply the actual language buyers use when they ask. And the Layer 1 technical baseline assesses whether AI systems can access and interpret tonic.ai at all, independent of what the content says. Nothing here is a finished conclusion — it is the set of assumptions the audit gets built on, and an assumption is far cheaper to fix now than a query set is to re-run later.
The validation call is a working session, not a presentation, and two kinds of decisions come out of it. The first is input validation: are the right entities in the right tiers, do the personas match who actually shows up in your deals, and are the capability ratings honest relative to the vendors you meet in live evaluations? Those answers set the audit architecture — query volume, competitive matchups, and the vocabulary every prompt is written in, across the AI platforms selected for this engagement. The second is engineering triage: several Layer 1 findings need no decision from anyone and can be in flight before we meet, so the audit measures an improved baseline rather than a known-broken one. The Pre-Call Checklist collects both lists in one place.
Three things worth knowing before you read the rest of this document.
What This Is This is the knowledge graph the audit will run against. Everything downstream — the buyer queries we generate, the competitors we test head-to-head, the language we phrase prompts in — is derived from the personas, competitors, capabilities, and pain points on these pages. It was built outside-in: from tonic.ai's own pages, competitor positioning, category listings, and review data, without access to your CRM or win/loss records. That is deliberate — it mirrors what an AI system can see about the unstructured de-identification and redaction category. It also means you know things we cannot.
What We Need From You Read the purple boxes. Each one names a specific entity we are uncertain about and states what changes in the audit if our reading is wrong. You do not need to review every card in detail — the purple questions are where your answer actually moves something. Everything else is here so you can check our work if you want to. The Pre-Call Checklist near the end collects every question in one printable list.
Confidence Badges Every entity carries a confidence badge. High means it came straight from a primary source — tonic.ai's own pages, a competitor's positioning, or a published review. Medium means we triangulated it from more than one indirect signal. Low means we inferred it and are asking you to confirm or kill it. Low-confidence items are not filler; they are the specific places where your judgment beats our research.
How AI systems would describe Tonic.ai from the outside — the naming, category, and positioning every generated query is anchored to.
→ Validate Is this audit measuring Tonic Textual, or the whole Tonic.ai platform? The seed URL was the Textual product page, so we scoped the category, the competitors, and the twelve capabilities to unstructured de-identification — which is why Delphix, K2View, and Synthesized are absent even though tonic.ai maintains dedicated /alternatives/ and /vs/ pages for all three. If the audit should cover Structural and Fabricate too, the test-data-management competitive set comes back in, the personas shift toward QA and platform engineering, and the query vocabulary changes from "redact PHI in clinical notes" to "database masking for test environments." One more wrinkle to settle in the same breath: the schema classifies Tonic.ai as a startup on headcount and funding, but the named customer list runs JPMorgan Chase, Qualcomm, Philips, Commonwealth Bank, and the NYC Department of Health — should query language read enterprise-evaluation even though the company is startup-segment?
6 personas — 3 decision-makers, 1 evaluator, 2 influencers — each of whom searches the unstructured de-identification category differently, which is what determines the intent and vocabulary of every query we generate.
Critical Review Area This is the section where your correction is worth the most. Personas drive query generation directly: each one produces a distinct cluster of prompts written at their altitude, in their vocabulary, with their evaluation criteria. A persona who does not exist in your deals burns query budget on the wrong buyer. A persona who exists but is missing here means an entire buying conversation goes unmeasured.
Data Sourcing Note Role, department, seniority, influence level, veto power, and technical level come straight from the knowledge graph and carry the confidence badge shown on each card. The role description, primary buying jobs, and query focus areas are synthesized — our reading of what that role does during a de-identification evaluation, based on the capabilities and pain points they map to. Names are placeholders for the role, not real people. Correct the synthesized lines freely; they are our inference, not your data.
→ Priya drives the evaluation but holds no veto in our model — if AI platform leads at your accounts actually own the spend, we promote her to decision-maker and add procurement-stage queries (pricing per 1,000 words, TCO against an in-house Presidio build) to her cluster.
→ We modeled Marcus as a hard veto on data leaving the VPC — if security instead signs off on a managed endpoint under a BAA, self-hosting stops being a gating query theme and Amazon Comprehend and Google Cloud DLP become live options inside his cluster rather than pre-eliminated ones.
→ Dana's query cluster assumes HIPAA expert determination is the deciding evidence — if GDPR and EU data residency is the more common blocker in your pipeline, we swap that set toward pseudonymization-under-GDPR language, which is a materially different vocabulary and surfaces a different competitive field.
→ Ethan is sourced from review mining at medium confidence and modeled as an influencer who inherits the evaluation — if data science actually initiates it, he becomes the entry-point persona and "our redaction script turned every document into a wall of [REDACTED]" moves to the top of the query set rather than the middle.
→ We gave Sofia veto power on the assumption VP Engineering owns the platform budget — if that money actually sits with the CISO or inside Priya's AI platform org, then one of our three veto-holders is wrong and the approval-stage queries are written for the wrong reader entirely.
→ Terrence is the only persona we inferred rather than sourced, and Guided Redaction is documented as beta — if records and legal redaction is not yet a funded buying motion, his entire query vocabulary gets reallocated to the AI platform and security personas rather than testing a market you are not selling into yet.
→ Who Else Shows Up? Three roles appear in de-identification deals often enough to be worth a check, and each would earn a dedicated query cluster if they show up in yours: a Chief Medical Information Officer or clinical data governance lead (if healthcare deals route through a clinical approval chain that is separate from the CISO's); a Data Protection Officer for EU entities (if GDPR is a distinct buying conversation from HIPAA rather than a subsection of it); and a third-party risk / vendor security reviewer (if the security questionnaire is its own gate with its own evaluation criteria, rather than something Marcus's team absorbs). Who else shows up in your deals?
11 competitors identified — 6 primary and 5 secondary — spanning named privacy vendors, hyperscaler APIs, and one open-source framework.
Why Tiers Matter Tier assignment decides where query budget goes. Each primary competitor gets a dedicated head-to-head set — roughly six to eight prompts apiece, so about 36–48 of the audit's queries — testing direct comparisons like "Tonic Textual vs Microsoft Presidio," "best alternative to Private AI for PHI redaction," and "Comprehend vs Tonic for clinical note de-identification." Secondary competitors are measured only for co-occurrence: do they show up when a buyer asks an open category question. Two of the six primaries — Protecto and Skyflow — carry medium confidence because we tiered them on category overlap and search co-occurrence rather than observed deal data. One more thing worth flagging before queries are written: Private AI rebranded to Limina in March 2026, models trained before then will overwhelmingly say "Private AI," and your own comparison page still lives at /vs/privateai-vs-tonic — we collapse both names to one entity so head-to-head share against your closest rival is not split in half and undercounted.
→ Validate Three questions, in order of how much they move the audit. First: in lost or stalled deals, how often is the real alternative "we'll just use Amazon Comprehend or Google DLP" versus a named privacy vendor like Limina or Protecto? Two of six primary slots currently go to hyperscaler APIs; if those are rarely serious contenders, we free roughly a dozen head-to-head queries and hand them to Limina and Protecto instead. Second: Protecto and Skyflow are tiered primary on category overlap and search co-occurrence, not on deal evidence — Protecto in particular is a heavy content marketer whose SEO footprint may exceed its real presence in your pipeline. If either rarely appears in an actual evaluation, moving them to secondary reallocates six to eight queries each. Third: is anyone here irrelevant, and is anyone missing? Unstructured is our leading removal candidate — a buyer may evaluate it as a complement to Textual Pipelines rather than an alternative to it — and any vendor that shows up in your deals but not on this page is currently invisible to the audit.
12 buyer-level capabilities — 7 strong, 4 moderate, 1 weak — rated outside-in, which is what determines which capability queries get tested and where the audit plays offense versus defense.
How reliably does it actually find every name, SSN, account number, and medical record ID buried in messy real-world documents — and what are its precision and recall numbers on data like mine?
Instead of blacking everything out, swap the real names and addresses for realistic fake ones so the document still reads naturally and my model still learns something from it.
Our sensitive data includes internal identifiers no off-the-shelf model has ever seen — can I teach it to catch those without hiring an ML team?
My sensitive data is in scanned PDFs, Word docs, PowerPoints, spreadsheets, JSON blobs and screenshots — will it handle all of that, or just plain text?
Our security team will not approve shipping regulated patient or customer text to a third-party API — can I run this entirely inside our own VPC or on-prem?
Does it plug into the S3 buckets, Databricks, Snowflake and Fabric we already run, and is there a real Python SDK and REST API so this can live in our pipeline instead of being a manual step?
When our auditor asks whether this data is genuinely de-identified under HIPAA or GDPR, I need documentation and a statistician's sign-off — not a vendor's marketing claim.
Can it transcribe and strip PII out of our contact center call recordings so we can finally use them for QA and model training?
Half our support tickets are in Spanish, German and Japanese — does detection quality hold up outside English or does it just over-redact everything?
A lawyer has to sign off on these redactions — I need a review screen, comments, and a timestamped record of who redacted what before anything goes out the door.
Turn our PDF and DOCX archive into clean, chunked markdown with the tables intact so we can embed it into a vector database without writing our own parser.
I need PII stripped inline, in milliseconds, on every prompt my customer-facing agent sends to the model — not in a batch job that runs overnight.
Which Three Lead? Seven of twelve capabilities are rated Strong: Sensitive Entity Detection Accuracy, Context-Preserving Synthetic Replacement, Custom Entity Types & Model Tuning, Document, Image & File Format Coverage, Self-Hosted Deployment & Data Residency, Cloud, Lakehouse & Developer Platform Integrations, and Compliance Evidence & HIPAA Expert Determination Support. The audit tests all 12, but competitive differentiation queries will emphasize 3. Which of these best represents where Tonic.ai actually wins deals — not where the product is objectively good, but where the buyer changes their mind?
→ Validate The ratings, then the gaps. The seven Strong ratings lean heavily on your October 2025 benchmark — aggregate F1 0.96 / precision 0.97 / recall 0.96 against Comprehend, Azure AI Language, Google Cloud DLP, Presidio, GLiNER, and spaCy — which is vendor-published, and we have read it as such. Two ratings we are least sure of: Audio & Call Recording De-Identification and Multilingual Detection Coverage are rated Moderate on relative rather than absolute grounds, because your comparison page asserts Limina lacks audio entirely while Limina's own site claims audio, DICOM, and 52 languages. One of those two public claims is stale — which languages are production-grade versus best-effort, and which audio formats are GA? And the single most consequential rating in this document: Real-Time Prompt & Response Redaction is the only Weak, because the Textual page promises to "secure real-time agentic workflows" but every path we could verify — datasets, pipelines, Python SDK, REST API, MCP server — is batch or request/response, not an inline low-latency proxy with policy enforcement. If a production runtime path exists with a published latency figure, tell us and we generate runtime head-to-heads against Skyflow, Protecto, and Nightfall; if not, we treat runtime as conceded and spend those queries on batch de-identification where you win. Finally: is any capability here really two capabilities, or two really one?
12 pain points — 6 high severity, 6 medium — written in the buyer's own words, because that phrasing is literally how the audit's queries get worded.
AI and product teams have a funded GenAI roadmap but cannot get legal or privacy approval to use real customer documents, clinical notes, or support transcripts for fine-tuning and RAG, so projects sit in review for months.
Regex and pattern-matching redaction tools black out so much of the text — including non-sensitive terms — that the resulting corpus loses the semantic context models need, making it useless for training or retrieval.
A single identifier missed during de-identification ends up embedded in a vector index or baked into model weights, where it cannot be surgically removed and can be surfaced by a user prompt — turning a recall gap into a reportable breach.
Teams that started with Presidio, spaCy, or hand-rolled regex now own an unfunded ML maintenance project — retraining recognizers, adding entity types, and building evaluation harnesses — instead of shipping product.
Managed cloud PII APIs require sending regulated text outside the organization's trust boundary, which fails security review, data residency rules, or BAA requirements — eliminating the cheapest options before evaluation even starts.
Analysts and paralegals redact documents by hand for FOIA responses, litigation production, and partner data sharing — weeks of effort per batch, with any single human miss becoming an unauthorized disclosure.
Sensitive values live inside scanned PDFs, images, slide decks, and email attachments where naive text extraction either fails outright or silently drops content, leaving PII undetected in the archive.
Contact center recordings and transcripts are the richest source of customer insight and support-agent training data, but sit unused because they are dense with names, card numbers, and account details.
Organizations cannot demonstrate to an auditor or regulator what was detected, what was transformed, who reviewed it, or what residual re-identification risk remains — so de-identification claims rest on vendor assertion rather than evidence.
Separate vendors handle database test data and document or free-text de-identification, producing two contracts, two security reviews, and inconsistent treatment of the same individual across structured and unstructured sources.
Batch de-identification cannot sit in the request path of a customer-facing agent, so live prompts and tool outputs reach the model unprotected while the offline corpus is carefully sanitized.
Detection quality degrades outside English, so global organizations either exclude non-English data from AI initiatives entirely or accept aggressive over-redaction that renders it unusable.
→ Validate Severity, phrasing, and one deliberate omission. Six of twelve pains are rated High, which means half this set competes for the same "test first" slot — if you had to name the three that actually stall deals, are they the legal-approval blocker, the Presidio maintenance burden, and the trust-boundary rule, or is that ordering wrong? On phrasing: the buyer language is what the audit's prompts get written in, so if "we traded a privacy problem for a garbage-data problem" is not how your buyers talk, tell us the sentence they actually say. And on omissions — we deliberately left out a pricing or TCO pain even though G2 reviews cite cost and setup complexity, because those complaints attach to Structural's database-masking workflow rather than Textual, which sells pay-as-you-go per 1,000 words. If Textual-specific pricing objections are real in your deals, that becomes a thirteenth pain point and a whole query cluster. Two others worth considering: quantifying residual re-identification risk (if buyers ask for a statistical risk number, not just a process), and throughput on very large archives (if "how long to process ten million documents" is a live objection rather than an afterthought).
50 pages of tonic.ai analyzed at the HTML level — crawler directives, sitemap, schema markup, heading structure, and date signals. 15 findings: 13 diagnostic, 2 requiring manual verification. Zero critical.
Engineering — Actionable Now There is no critical blocker here: tonic.ai serves complete server-rendered HTML on all 50 pages, robots.txt permits every declared AI crawler, and schema coverage sits at 0.84. The four high-severity items are all engineering-owned and all shippable in under a day each. Start with the Cloudflare check — every response carries server: cloudflare, and Cloudflare's Bot Fight Mode and "Block AI bots" managed rules operate at the edge, above robots.txt. Cloudflare changed its default to block AI crawlers in July 2025 and its network handles roughly 24% of all web traffic (Cloudflare, July 2025), so a permissive robots.txt is not proof of access; confirm from logs that GPTBot, ClaudeBot, PerplexityBot, and Google-Extended got 200s over the last 30 days. If they did not, that supersedes everything else in this document. Then: remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template, which currently tells a reader that the Limina and K2View pages are 2023 comparisons of Delphix; and emit <lastmod> from the Webflow CMS across all 1,878 sitemap URLs, which today publish none. The two Content-owned findings below are noted for planning, not for this sprint.
What we found: All three /alternatives/ pages emit a JSON-LD dateModified of 2026-07-16, but the visible body text says something different on each: /alternatives/limina reads "Last updated August 2025", /alternatives/k2view reads "Last updated December 2025", /alternatives/delphix reads "Last updated July 2026". Worse, all three also render a second, orphaned block reading "Book a DemoLast updated December 2023. Comparison based on Delphix's full suite of services." — so the Limina page and the K2View page both tell a reader that they are a December 2023 comparison of Delphix. The /vs/ pages are clean (single "Last updated July 2026" stamp).
Why it matters: Comparison pages are the single most-cited page type in vendor-evaluation queries, and an LLM reads the visible text, not the JSON-LD. On the Limina page — the head-to-head against the closest competitor in the knowledge graph — the strongest recency signal a model can extract is "December 2023", and the strongest topical signal in that same sentence names the wrong competitor. That is both a freshness penalty and an accuracy defect on the highest-value asset in the competitive set.
Recommended fix: Remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template (it is clearly an un-cleared static element in the component). Then bind the visible "Last updated" string to the same CMS field that feeds JSON-LD dateModified so the two can never diverge, and re-stamp Limina and K2View to their true review dates.
What we found: The Synthesized comparison exists at both /vs/synthesized (2,460 words) and /blog/tonic-vs-synthesized (2,450 words) with the same H1 and near-identical body, and each page self-canonicalizes. Separately, /alternatives/delphix sets rel=canonical to /vs/delphix, yet both URLs are submitted in sitemap.xml as independent entries, and a third Delphix comparison is reachable from /vs/delphix's own "Full comparison article" link. The Delphix comparison therefore occupies three URLs and the Synthesized comparison two.
Why it matters: Splitting one comparison across two self-canonical URLs divides link equity and retrieval signals between them, and gives retrieval-augmented systems two competing candidates for the same query with no signal about which is authoritative. Submitting a URL in the sitemap that canonicalizes elsewhere sends contradictory instructions to crawlers and wastes crawl budget on a page the site itself says is not the original.
Recommended fix: Pick one canonical home per competitor — the /vs/{competitor} path is the stronger pattern — and 301 the duplicates to it, or at minimum set rel=canonical on /blog/tonic-vs-synthesized to /vs/synthesized. Remove any URL that canonicalizes elsewhere (currently /alternatives/delphix) from sitemap.xml.
What we found: https://www.tonic.ai/sitemap.xml is a flat (non-index) sitemap listing 1,878 URLs. Every entry contains only a <loc> element — zero URLs carry <lastmod>, and zero carry <priority> or <changefreq>. The HTTP Last-Modified header is no substitute: the origin returned last-modified: Mon, 27 Jul 2026 16:36:02 GMT — the moment of our request — for pages whose content is years old, so it is a build timestamp, not a content timestamp.
Why it matters: With no <lastmod> and a header that always reports "now", a crawler has no machine-readable way to tell which of 1,878 pages changed. Recrawl scheduling falls back to heuristics, so genuinely refreshed pages — including the comparison pages Tonic actively maintains — wait in the same queue as a 2021 careers post. Freshness is a documented driver of AI citation selection, and the site currently forfeits the cheapest signal available to it.
Recommended fix: Emit <lastmod> from the Webflow CMS "last published" field for every URL in sitemap.xml, and stop returning a request-time Last-Modified header on cached HTML (serve the origin's real content timestamp, or omit the header and rely on ETag).
What we found: Of the 1,878 URLs in sitemap.xml, 1,443 (76.8%) are per-release changelog entries: 820 under /tonic-structural-release-notes, 447 under /tonic-textual-release-notes, 172 under /tonic-fabricate-release-notes, and 4 under /product-release-notes. Only 435 URLs are commercial or editorial content. The sitemap is a single flat file with no index and no segmentation.
Why it matters: Release-note pages are near-duplicates of one another, carry almost no standalone text, and are essentially never the right answer to a buyer question — yet they outnumber the product, comparison, capability, and case-study pages by more than three to one in the only prioritization signal the site publishes. Crawl budget spent re-fetching 1,443 changelog entries is crawl budget not spent on the benchmark and comparison pages that actually win citations.
Recommended fix: Convert sitemap.xml into a sitemap index with separate child sitemaps (pages, blog, guides, case studies, release notes) so crawlers can prioritize, and consider consolidating per-release entries into one paginated changelog per product rather than one indexable URL per release.
What we found: Every /blog/ post and every /case-study/ page we inspected emits datePublished in its BlogPosting/Article JSON-LD but no dateModified — 11 of 11 blog and case-study pages in the inventory. The /vs/, /alternatives/, /ai-model-benchmarks/ and guide templates do emit both. The result is that the only machine-readable date on a blog post is the day it first shipped, and pages like /blog/redacting-sensitive-free-text-data-build-vs-buy (published 2023-12-12) and /blog/how-to-de-identify-legal-documents-with-tonic-textual (2024-04-15) present as 958 and 833 days old respectively even if they have been edited since.
Why it matters: Adding dateModified is the difference between a refresh being invisible and a refresh being credited. Without it, the content team can rewrite a post and gain nothing in freshness-weighted retrieval, which removes the incentive to maintain the archive at all. It also makes the stale-content problem below impossible to fix without republishing under a new date.
Recommended fix: Add a dateModified property to the BlogPosting/Article JSON-LD in the blog and case-study templates, bound to the CMS "last updated" field, and surface the same value as a visible "Updated {date}" line beneath the byline.
What we found: Dated content-marketing pages in the inventory that fall outside the citation window: /blog/redacting-sensitive-free-text-data-build-vs-buy (2023-12-12, 958 days), /blog/how-to-de-identify-legal-documents-with-tonic-textual (2024-04-15, 833 days), /case-study/tax-agency-unblocked-ai-llm-workflows-with-secure-data-synthesis (2025-05-02, 451 days), /case-study/ontra-accelerates-ai-innovation-in-legal-tech-with-tonic-textual (2025-11-11, 258 days), and /blog/self-serve-custom-entity-types-in-tonic-textual (2025-11-17, 252 days). Separately, the three /playbooks/ pages publish no date at all in either the rendered body or JSON-LD. The two >365-day posts are substantive (content depth 0.80 and 0.75) and are the site's primary build-vs-buy and document-redaction walkthroughs.
Why it matters: These are not low-value pages being allowed to age — the build-vs-buy post is the natural landing point for the "we built our own PII scrubber on Presidio" buyer, and the legal-documents walkthrough is the only step-by-step de-identification tutorial on the site. Both now describe a product that has since gained Custom Entity Types, audio, an MCP server, and LLM-assisted NER, so they are simultaneously the stalest and the least accurate pages in the set.
Recommended fix: Refresh the two >365-day posts against current Textual capability and republish with dateModified set (see the schema finding above). Add a visible publish/updated date to the three /playbooks/ pages, which currently carry no date signal of any kind.
What we found: Every page in the inventory has exactly one H1, which is good. The problems are below it. /integrations nests 96 H3 headings under just 2 H2s, so the connector directory has effectively no hierarchy. The three /playbooks/ pages use 6 H2s and zero H3s, leaving a 35-entry entity-type glossary — roughly half the page's word count — completely unheaded. /products/textual uses a full two-sentence marketing paragraph ("Instantly redact sensitive PII from text and audio to safely train models, and secure real-time agentic workflows. Get the essential privacy layer for every LLM interaction.") as an H2. /pricing repeats the generic headings "Features", "Common questions", and "Enterprise" three times each, once per product block, with nothing distinguishing them. The homepage carries 94 H3s, nearly all of which are resource-card titles rather than page structure.
Why it matters: Headings are how a retrieval system segments a page into citable passages and how it labels what each passage is about. A 96-item flat list, an unheaded glossary, a sentence used as a heading, and three identical "Features" headings all produce passages that either cannot be delimited or cannot be told apart, which pushes otherwise useful content below the citation threshold.
Recommended fix: Group /integrations connectors under category H2s (Relational, NoSQL, Warehouse/Lakehouse, Files, SaaS) with connector names as H3. Give the playbook entity glossary an H2 and group entities by class. Replace the /products/textual paragraph-H2 with a short noun phrase and move the sentence into body copy. Qualify the repeated /pricing headings ("Fabricate features", "Structural common questions", and so on).
What we found: Schema coverage across the site is generally strong — 33 of 50 inventoried pages carry an appropriate specific type (Product, SoftwareApplication, FAQPage, TechArticle, Review, HowTo, CollectionPage). Nine do not: all four /partners/ pages, all three /playbooks/ pages, /integrations, and /products/validate emit only the global Organization / WebSite / Person / PostalAddress / ContactPoint block that appears in the site footer template, with no WebPage, Product, HowTo, or BreadcrumbList of their own. Separately, docs.tonic.ai/textual emits no JSON-LD at all and has no meta description.
Why it matters: Those nine pages carry the marketplace, lakehouse, and hyperscaler integration story — exactly the material a buyer query like "PII redaction inside Snowflake" should retrieve — and the playbooks are the only procedural how-to content on the marketing site. Without HowTo, Product, or even BreadcrumbList markup, they are unlabeled prose in a corpus where their siblings are richly typed, and they lose the breadcrumb trail that helps a model place them in the site hierarchy.
Recommended fix: Add BreadcrumbList and WebPage to the /partners/, /playbooks/, /integrations and /products/validate templates; add HowTo to the playbook template (each already has a numbered "Playbook steps" section that maps directly onto HowToStep); add SoftwareApplication to /products/validate. On docs.tonic.ai, add a meta description and TechArticle markup to the GitBook theme.
What we found: Body word counts after stripping nav and footer: /products/validate 255 words (content depth 0.30), /partners/aws 319, /partners/databricks 325, /partners/microsoft/fabric 383, /products/tonic-datasets 444, /partners/snowflake/textual-native-app 459 (all 0.40–0.50), plus /integrations at 1,684 words that are almost entirely one-sentence connector blurbs (0.45) and /playbooks/audio-redaction-and-synthesis where roughly 380 of 623 words are a repeated entity-type glossary shared verbatim with two other playbooks (0.40). For comparison, the same site's /ai-model-benchmarks/textual-benchmark page runs 6,424 words with per-competitor precision/recall/F1 tables.
Why it matters: Every one of these thin pages sits on a query a buyer actually asks — "Tonic Textual on Databricks", "redact PII in Snowflake", "de-identify audio recordings". A 320-word page of two-sentence claims gives a retrieval system nothing specific enough to quote, so the query is answered from a competitor's or a hyperscaler's documentation instead. The gap is especially stark because Tonic clearly knows how to write the deep version — it did so for the benchmarks.
Recommended fix: Bring each partner page up to a concrete implementation narrative: the specific integration mechanism, a worked example, throughput or accuracy figures where they exist, and the setup steps. Replace the shared entity glossary on the playbooks with a single canonical entity-types reference page and link to it. Expand /integrations connector entries with supported versions, file types, and limits.
What we found: The Resources module renders the literal string "No items found." in the served HTML of /alternatives/delphix, /alternatives/k2view, /capabilities/expert-determination, /capabilities/government-redaction, and /case-study/ontra-accelerates-ai-innovation-in-legal-tech-with-tonic-textual. On /alternatives/delphix the module heading promises "Further reading and resources on Delphix alternatives" and is followed immediately by the empty-state string.
Why it matters: This is text a crawler ingests as page content. It adds an empty-state string to pages that are otherwise arguing a competitive case, and it removes the internal links that would normally pass authority from a comparison page into the supporting benchmark and blog content. On the expert-determination and government-redaction pages it also strips the only outbound paths to the deeper compliance material.
Recommended fix: Populate the Resources collection on these five pages with existing relevant content — the Textual benchmark and PrivacyBench pages for the comparison pages, the HIPAA expert-determination blog posts for /capabilities/expert-determination, the FOIA redaction post for /capabilities/government-redaction — and change the Webflow component to hide the whole module when the collection is empty rather than printing the placeholder.
What we found: /partners/microsoft/fabric renders "Choose an output OneLake folder for storing the de-identified files." three times consecutively and "If needed, adjust settings and reprocess." four times consecutively inside its numbered setup sequence — on a page that only has 383 words of body copy, roughly 10% of the page is a duplicated sentence. /solutions/use-case/llm-privacy-proxy displays its headline metric as "<0 min — Real-time redaction", which reads as "less than zero minutes". /ai-training-data renders its "On this page" block and "Generate the AI training data you need" CTA twice in a row.
Why it matters: These are small but they land on the two pages most relevant to the knowledge graph's weakest feature area. "<0 min" is the only latency claim anywhere on the site for real-time redaction, and as rendered it is not a number a model can quote or a buyer can trust. The Fabric repetition makes the shortest integration page look automated rather than authored.
Recommended fix: Remove the duplicated CMS repeater rows on /partners/microsoft/fabric and /ai-training-data. Replace the "<0 min" stat on /solutions/use-case/llm-privacy-proxy with a real, defensible latency figure and its measurement conditions, or remove the stat block.
What we found: og:title and og:description are present sitewide, but og:image is absent on /faqs, /vs/delphix, /vs/synthesized, /partners/aws, /partners/databricks, /partners/microsoft/fabric, and /partners/snowflake/textual-native-app. Meta descriptions are present on all 49 www pages inspected (93–254 characters) and missing only on docs.tonic.ai/textual.
Why it matters: Low impact on retrieval itself, but /vs/delphix and /vs/synthesized are the pages most likely to be shared into Slack, LinkedIn, and sales email during an active evaluation, and they currently unfurl without a preview image. It is a cheap fix on the pages where social amplification matters most.
Recommended fix: Add an og:image to the /vs/, /partners/, and /faqs templates, defaulting to the existing Tonic.ai social card where no bespoke image exists. Add a meta description to docs.tonic.ai/textual.
What we found: The homepage contains href="https://textual.tonic.ai/signup](https://textual.tonic.ai/signup" — a markdown link fragment ("](") pasted into a Webflow link field, producing a single URL with the closing-bracket sequence embedded. This is the only malformed href we found across the 50 pages fetched, and a 40-URL random sample of non-release-note sitemap entries returned HTTP 200 for all 40, so there is no broader broken-link problem.
Why it matters: Minor on its own — one dead signup path on the highest-traffic page — but worth fixing because it is a conversion link and because it indicates content is being pasted from markdown into the CMS without validation, which is the same class of error that produced the orphaned "Last updated December 2023" block on the comparison pages.
Recommended fix: Correct the link to https://textual.tonic.ai/signup, and add a build-time link validation step (or a Webflow pre-publish check) that rejects hrefs containing "](" or whitespace.
The following items could not be assessed through our analysis method (rendered markdown). We recommend your engineering team verify these manually before the validation call.
What to check: https://www.tonic.ai/robots.txt contains a single User-agent: * group that disallows only /styleguide, so GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Googlebot and Bytespider are all permitted. docs.tonic.ai/robots.txt goes further and publishes Content-Signal: ai-train=yes, search=yes, ai-input=yes with Allow: /. However, every response from www.tonic.ai carries server: cloudflare, and Cloudflare's Bot Fight Mode, AI Labyrinth, and "Block AI Bots" managed rules operate at the edge, independent of robots.txt — a site can publish a fully permissive robots.txt and still return 403s to GPTBot and ClaudeBot. We fetched with a browser user-agent and cannot determine from the outside how the edge treats declared AI user-agents. This is the single highest-leverage technical check on the site: if the WAF is silently blocking AI crawlers, nothing downstream can produce citations, and confirming 200s converts this from a risk into a verified strength worth stating in the audit.
Recommended action: In the Cloudflare dashboard, check Security > Bots for Bot Fight Mode, the "Block AI bots" managed rule, and any custom WAF/rate-limit rules matching AI user-agents; then confirm against origin/Cloudflare logs that GPTBot, ClaudeBot, PerplexityBot and Google-Extended requests returned 200 over the last 30 days. Explicitly allowlist the seven crawlers if any rule intercepts them.
What to check: This analysis parsed the raw server-rendered HTML returned by the origin, which let us assess schema markup, meta tags, OG tags, and content directly rather than inferring them. All 50 pages returned full body content, headings, and JSON-LD without executing JavaScript, so no page shows client-side-rendering symptoms. What we could not check is the opposite direction: whether the browser-rendered DOM adds or reorders content the server HTML does not contain. The homepage's four-panel tabbed hero, the /integrations and /customers filter widgets, and the FAQ accordions are interactive components whose full text we did find in the server HTML, but whose behaviour under a JS-executing crawler we did not observe. Server-rendered HTML is the stronger position to be in and this is a confirmation step, not a suspected defect — but a mismatch between server HTML and rendered DOM is the one failure mode our method structurally cannot see.
Recommended action: Spot-check five pages (homepage, /products/textual, /integrations, /pricing, /vs/synthesized) in Google Search Console's URL Inspection "View crawled page" and in a headless-Chrome crawler, comparing rendered text and JSON-LD against view-source. Confirm the homepage hero's non-active tab panels and the /integrations connector list are present without JavaScript.
Partial Sample This analysis covers 50 pages of the 435 commercially relevant URLs in sitemap.xml — roughly 12%, weighted toward product, comparison, capability, partner, and case-study pages. Sitewide claims about schema, headings, and thin content are extrapolated from that sample and should be read as representative, not exhaustive. Freshness is a bigger caveat: 26 of the 50 pages could not be scored at all because they publish no date in either the rendered body or JSON-LD — 23 product and commercial pages and all 3 structural/reference pages — so the 0.606 weighted score reflects only the 24 dated content-marketing pages. Fixing the <lastmod> and dateModified findings above would make the next measurement materially more complete.
Why Now
Once the inputs on these pages are confirmed, the audit measures how Tonic.ai actually appears across the selected AI platforms for the questions your buyers ask — "how do I de-identify clinical notes for LLM training," "best alternative to Private AI for PHI redaction," "build vs buy a PII scrubber on Presidio," "can I run PII redaction inside my own VPC." You will see exactly which of those queries return answers naming Limina, Presidio, Comprehend, or Skyflow but not Tonic.ai, which name you but describe you wrongly, and what specifically it would take to appear in the ones you are missing. Fixing the Layer 1 items in the meantime means the audit measures a repaired baseline rather than a known-broken one — the Cloudflare check in particular determines whether any of the rest of it is measurable at all.
45–60 minutes. We walk through this document together, resolve the questions in the Pre-Call Checklist, and lock the competitor tiers, personas, and capability ratings the query set will be built from.
We generate buyer queries from the validated personas, pain points, and capabilities, then run them across the AI platforms selected for this engagement and capture every response and citation.
Visibility analysis, competitive positioning against the confirmed set, and a three-layer action plan — technical, content, and authority — prioritized by which gaps actually cost you citations.
Engineering Can Start Today Three Layer 1 fixes need no decision from anyone at the validation call. One: verify Cloudflare is not intercepting AI crawlers above robots.txt — check Security > Bots for Bot Fight Mode and the "Block AI bots" managed rule, then confirm from logs that GPTBot, ClaudeBot, PerplexityBot, and Google-Extended received 200s over the last 30 days. Crawler access is currently unverified, and if the edge is blocking them it supersedes everything else in this document. Two: remove the orphaned "Book a Demo / Last updated December 2023 / Delphix" block from the /alternatives/ Webflow template and bind the visible "Last updated" string to the same CMS field that feeds JSON-LD dateModified — a sub-day fix on the Limina page, which is the head-to-head against your closest competitor. Three: emit <lastmod> from the Webflow CMS across all 1,878 sitemap URLs and stop returning a request-time Last-Modified header. These don't depend on the rest of the audit and will improve your baseline visibility before we even measure it.
Two jobs before we meet. The questions on the left require your judgment — no one knows your business better than you. The engineering tasks on the right don't require the call at all.