Engagement Foundation Review

Graylog Audit Foundation

Before we run the audit, we need to make sure we're asking the right questions about the right competitors to the right buyers. This document presents what we've learned about Graylog's market — your job is to tell us what we got right, what we got wrong, and what we missed.

Prepared July 23, 2026
graylog.org
Log Management & SIEM
GEO Readiness

Where You Stand Today

Before we measure Graylog's citation visibility in the log management and SIEM space, these three signals tell us whether AI crawlers can access and trust the site at all. Right now all three read red — a crawler-access problem sits upstream of every content decision.

Technical Readiness
At Risk
One critical finding: a WAF/bot-management layer returns HTTP 403 to non-browser clients site-wide, so AI crawlers are served a block instead of content. Two high-severity findings compound it — an unreachable robots.txt/sitemap.xml (403) and no page-level freshness signals.
Content Freshness
At Risk
Weighted freshness 0.20. No page on the site carries a detectable published or updated date, and the sitemap that would supply lastmod is 403. All 7 content-marketing pages (6 competitor comparison pages + the "Mastering SIEM" guide) score 0.20 — none updated within 90 days, all older than 180 days. 28 product/commercial pages have no detectable date — verify manually. AI-cited content runs 25.7% fresher on average than Google organic results (Ahrefs, August 2025).
Crawl Coverage
At Risk
robots.txt and sitemap.xml both return HTTP 403 to automated agents, and the WAF escalates to an IP-level 403 across all paths after ~11 requests — crawler access is effectively blocked and unverified. Crawl directives are unreadable and no URL/lastmod inventory is available. This regressed since the 2026-06-14 crawl, when both files were publicly retrievable.
Executive Summary

What You Need to Know

AI search is reshaping how log management and SIEM buyers discover and evaluate solutions — lean security and IT operations teams increasingly open an AI chatbot before they open a browser tab. Companies that establish citation visibility in this shift now gain a first-mover advantage that compounds: early citations become self-reinforcing as AI platforms learn which domains to trust. Graylog enters this moment as a well-defined mid-market challenger with a clear cost-and-simplicity story — exactly the kind of position AI answers can amplify, or overlook.

This Foundation Review presents the inputs that will drive Graylog's audit, so we can validate them together before any query runs. It covers the competitive landscape that shapes how we construct head-to-head queries, the buyer personas that determine which search intents we test, and the Layer 1 technical baseline that determines whether AI platforms can reach Graylog's content in the first place. Each section is something we're asking you to confirm or correct — not a finished conclusion.

The validation call is a working session with real stakes: your answers decide which inputs drive the buyer query set across the AI platforms we test. Two kinds of decisions come out of it — input validation (are the right entities in the right tiers, and is anyone inferred rather than sourced?) and engineering triage (which technical fixes can start before results come back?). The cards below carry the specifics; the pre-call checklist aggregates every decision into one page.

TL;DR — Action Items
  • 🔴 Critical: WAF/bot protection returns HTTP 403 to automated crawlers site-wide — engineering should allow-list the major AI crawler user-agents (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) and their IP ranges so they receive 200s instead of a block; this can zero out AI visibility upstream of everything else.
  • 🟡 High: robots.txt and sitemap.xml return 403 — restore both to a 200 for every user-agent so crawl directives and the lastmod inventory become readable again.
  • 🟣 Validate at the Call: Susan Albright (IT Compliance / GRC Manager) — this persona is inferred, not sourced; if a dedicated GRC buyer doesn't actually sit in Graylog's committee, we drop an entire compliance-evidence query cluster from the audit.
  • 🟣 Validate at the Call: Wazuh's primary tier — if free, open-source Wazuh shows up mostly in DIY / build-vs-buy conversations rather than head-to-head deals, moving it to secondary shifts ~5 direct-comparison queries out of the head-to-head set.
  • ✅ Start Now: Restore public sitemap.xml with accurate lastmod — engineering can exempt robots.txt and sitemap.xml from WAF challenges immediately; it doesn't depend on any validation-call decision.
  • 📋 Validation Call: Confirm the WAF now serves AI crawlers 200s before queries run — until crawlers can fetch graylog.org, the audit would measure a blocked site rather than Graylog's real visibility.
How This Works

Reading This Document

The Point This document is the foundation for Graylog's GEO audit. Before we test how AI engines answer log management and SIEM buying questions, we validate the inputs that drive those tests — the competitors, buyer personas, capabilities, and pain points below. Get these right and the audit measures what matters; get them wrong and we test the wrong queries.

Your Job Tell us what we got right, what we got wrong, and what we missed. The purple boxes throughout flag the specific items where your answer changes how the audit is built — those are the questions to come to the call ready to answer.

Confidence Badges Each item carries a High / Medium / Low badge showing how firmly it's sourced. High means it came straight from Graylog's site or reviewer language. Medium means we inferred it and it's worth confirming. Low means it's a genuine open question for the call.

Company Profile

Who We're Auditing

Company Graylog High
Domain graylog.org
Name variants Graylog Inc · Graylog Security · Graylog Enterprise · Graylog Open · GrayLog · Gray Log
Category Log management & SIEM with API security monitoring
Segment Mid-market
Key products Graylog Security · Enterprise · Open · API Security · Cloud
Positioning Centralized log collection, threat detection, and security analytics purpose-built for lean security and IT operations teams — predictable, volume-based pricing versus ingestion-priced incumbents.

→ Confirm Graylog's category bundles three buying conversations — centralized log management, SIEM threat detection, and API security monitoring — across an open-core model (free Graylog Open plus paid Security/Enterprise/Cloud). Is this one unified evaluation, or does API Security in particular run as a separate buying motion with a different buyer? If it's separate, we split it into its own query cluster instead of testing it inside core SIEM queries; if it's unified, we keep one cluster and save the query budget.

Buyer Personas

Who Buys Graylog

5 personas — 1 decision-maker, 1 evaluator, 3 influencers. Personas drive the query set: each one searches a mid-market SIEM purchase differently, so who's really in the committee determines which intents we test.

Critical Review Area Personas are the highest-leverage input in this document. If we have the buying committee wrong — a missing role, a misjudged influence level, an inferred persona who isn't really there — the audit tests the wrong search intents. Scrutinize these harder than anything else.

Data Sourcing Note Role, seniority, department, influence, and veto power are pulled from the KG (most from G2/review mining). The buying jobs and query focus areas on each card are synthesized from those fields plus the pain points each role owns — they're our read of how each persona searches, and exactly the kind of thing to correct.

Marcus Bell
CISO / VP of Security · C-Suite
Decision-maker High
Owns the security program and the SIEM budget line for a lean mid-market security org; signs the contract and answers to the board on risk and spend.
Veto power: Yes — final sign-off on the log/SIEM platform purchase.
Technical level: Medium.
Primary buying jobs: Sets evaluation criteria, weighs total cost of ownership against risk reduction, approves the shortlist and the final vendor.
Query focus areas: Cost predictability vs. Splunk-style overages, risk and compliance coverage, "best SIEM for lean security teams," vendor consolidation.
Source: G2 / review mining

In a 3–5 person security org, does Marcus personally run the technical evaluation, or delegate it to Elena? If he's hands-on, executive and practitioner queries collapse into one cluster instead of two.

Elena Vasquez
Security Operations Manager / SOC Lead · Manager
Evaluator High
Runs day-to-day security operations and lives inside the tool; the person who judges whether a SIEM is actually operable by a lean team.
Veto power: No (per KG) — high influence, drives the hands-on evaluation and recommendation.
Technical level: High.
Primary buying jobs: Runs the proof-of-concept, judges alert quality and operability, recommends up to Marcus.
Query focus areas: Day-to-day operability, alert tuning and fatigue, detection coverage, "SIEM a small team can actually run."
Source: G2 / review mining

Elena is tagged evaluator with no formal veto — but in a lean SOC she may be the one who quietly kills or advances a tool. Does she hold practical veto? If yes, we reclassify her as a decision-maker and pull day-to-day operability queries into the decision set.

David Okonkwo
Director of IT Operations / Infrastructure · Director
Influencer High
Owns the infrastructure the platform runs on and the logging pipeline that feeds it; often the one feeling the cost-and-scale pain first.
Veto power: No — medium influence on deployment fit and run cost.
Technical level: High.
Primary buying jobs: Assesses infrastructure fit, deployment and day-2 ops burden, and total run cost; weighs self-hosted vs. SaaS.
Query focus areas: Log infrastructure cost and scale, deployment/ops effort, "self-hosted vs. SaaS SIEM," storage and retention.
Source: G2 / review mining

The cost-explosion pain lands partly on IT Ops. Does David control the infrastructure/logging budget line, or only influence it? If he owns budget, he moves to decision-maker and cost-predictability queries gain a second buyer to satisfy.

Priya Raman
Senior Security / Detection Engineer · Senior IC
Influencer High
The hands-on practitioner who writes detections, builds parsers, and searches logs during incidents; often the first person to touch the product.
Veto power: No — medium influence, strong technical credibility with the committee.
Technical level: High.
Primary buying jobs: Trials the product, stress-tests detection content, parsing, and search speed; champions (or blocks) bottom-up.
Query focus areas: Detection content and parsing, search speed during incidents, open-source flexibility, "Graylog vs. Elastic for detection engineering."
Source: G2 / review mining

Graylog Open is free and open-source, so engineers like Priya often bring it in bottom-up. Does she initiate evaluations, or only execute a decision made above her? If she's the entry point, we add practitioner-stage queries (detection content, parsing, search speed) that a top-down persona set would miss.

Susan Albright
IT Compliance / GRC Manager · Manager
Influencer Medium
Responsible for proving log retention and producing audit evidence (PCI, HIPAA, SOC 2); cares about the platform as a compliance instrument, not a threat-hunting tool.
Veto power: No — medium influence on compliance-driven requirements.
Technical level: Low.
Primary buying jobs: Checks compliance and retention coverage, validates audit/reporting fit against required frameworks.
Query focus areas: Compliance evidence and retention proof, "SIEM for PCI / HIPAA / SOC 2," scheduled audit reporting.
Source: LLM inference (not directly sourced)

Susan is inferred, not sourced — we're not certain a dedicated GRC/compliance buyer sits in Graylog's committee versus compliance being owned by Marcus. Does a compliance manager actually evaluate the tool? If not, we drop this persona and its PCI/HIPAA/SOC 2 query cluster; if yes, we keep it and give those queries real weight.

Who Else Shows Up? These roles sometimes appear in log/SIEM deals — do they show up in yours? MSP / MSSP partner buyer (if Graylog sells meaningfully through partners who resell or operate it, they search as a distinct buyer); DevOps / Platform Engineering lead (owns log pipelines and observability, may drive the log-management side of the decision); Procurement / Finance (in cost-driven, Splunk-replacement deals, a budget owner can shape the shortlist). Who else is in the room when these deals close?

Competitive Landscape

Who You're Up Against

5 primary + 5 secondary competitors identified. Tier assignments decide which vendors get head-to-head query coverage in the log management and SIEM audit.

Why Tiers Matter Each primary competitor gets roughly five or more head-to-head queries — around 25–40 direct-comparison queries across the five primaries. Queries like "Graylog vs. Splunk," "best Splunk alternative with predictable pricing," and "open-source SIEM for lean teams" only fire against vendors we tier primary. We're least certain about Wazuh: it's tiered primary, but if free, open-source Wazuh shows up mainly in DIY / build-vs-buy conversations rather than head-to-head deals, moving it to secondary would shift roughly five queries out of the direct-comparison set.

Primary Competitors

Splunk

PrimaryHigh
splunk.com
The enterprise SIEM and log analytics incumbent (now a Cisco company), extremely powerful and extensible but priced by ingestion volume — the cost-explosion pain that Graylog explicitly positions against with predictable, volume-based licensing.
Source: automated scrape

Elastic Security

PrimaryHigh
elastic.co
Search-first open-source stack (Elasticsearch/Kibana) with a security layer; strong on visualization and petabyte-scale search but requires significant tuning and ops effort, and buyers often weigh it directly against Graylog as the other open-core log platform.
Source: category listing

Wazuh

PrimaryHigh
wazuh.com
Free, fully open-source XDR/SIEM with agent-based endpoint telemetry; the closest budget-driven alternative for lean teams, though it lacks Graylog's polished search, enterprise support, and API security coverage. Tier flagged for confirmation — see the question below.
Source: category listing

Microsoft Sentinel

PrimaryHigh
microsoft.com
Cloud-native SIEM tightly integrated with the Microsoft/Azure ecosystem; strong for Microsoft-centric shops with built-in automation, but consumption-based pricing and Azure lock-in are friction points Graylog counters on cost predictability and deployment flexibility.
Source: category listing

Sumo Logic

PrimaryHigh
sumologic.com
Cloud-native log analytics and SIEM aimed at larger organizations; strong real-time analytics and managed SaaS delivery, but higher cost at scale and a SaaS-only model make it a direct comparison for teams weighing Graylog's self-hosted and hybrid options.
Source: category listing

Secondary Competitors

Datadog

SecondaryMedium
datadoghq.com
Observability-first SaaS platform that has extended into log management and Cloud SIEM; wins DevOps-led buyers but ingestion/indexing pricing gets expensive fast, and security depth is secondary to monitoring — an adjacent overlap rather than a core SIEM head-to-head.
Source: category listing

Exabeam

SecondaryMedium
exabeam.com
Security-analytics-led SIEM built around UEBA and behavioral modeling (now merged with LogRhythm); stronger prebuilt detection and analytics than Graylog but heavier, costlier, and oriented to larger SOCs rather than lean mid-market teams.
Source: category listing

Devo

SecondaryMedium
devo.com
Cloud-native security data platform positioned on fast querying of large log volumes and hot data retention; appears in analyst comparisons with Graylog but is enterprise-priced and less commonly evaluated by cost-sensitive mid-market buyers.
Source: category listing

ManageEngine Log360

SecondaryMedium
manageengine.com
Bundled SIEM and log management from the ManageEngine/Zoho IT suite with cloud and on-prem options; appeals to IT-ops-led buyers who want compliance reporting out of the box, but is less flexible on custom parsing and large-scale search than Graylog.
Source: category listing

Grafana Loki

SecondaryMedium
grafana.com
Lightweight, cost-efficient open-source log aggregation paired with Grafana dashboards; excellent visualization and cheap storage but a log/observability tool without native SIEM threat detection, so it overlaps on log management only.
Source: category listing

→ Confirm the Set Three things to check: (1) Missing vendors — do IBM QRadar, Cribl, or CrowdStrike Falcon LogScale (Humio) show up in your deals? Any of them could warrant a primary slot. (2) Wazuh's tier — is free, open-source Wazuh really a head-to-head primary, or does it live in DIY / build-vs-buy conversations? Moving it to secondary reallocates ~5 head-to-head queries. (3) Irrelevant listings — do Grafana Loki or ManageEngine Log360 actually surface in your deals, or are they category noise we should drop so we don't spend queries on them?

Feature Taxonomy

Capabilities We'll Test

12 buyer-level capabilities mapped — 4 rated Strong, 7 Moderate, 1 Weak. These determine which capability queries the audit runs and which get emphasized in competitive differentiation.

Centralized Log Management & Collection Strong High

Collect, ingest, and centralize logs from every source — servers, network gear, cloud, and apps — into one searchable place.

Fast Search & Incident Investigation Strong High

Search across terabytes of logs in seconds to run down an incident before it spreads.

Log Parsing, Pipelines & Data Routing Strong High

Normalize messy logs from dozens of formats and route or drop data with rules before it hits storage.

Predictable, Volume-Based Pricing Strong High

Know what my SIEM will cost next year — no surprise ingestion overages that price me out of my own logs.

Threat Detection & SIEM Correlation Moderate Medium

Detect real threats with correlation rules and prebuilt detection content instead of writing everything from scratch.

UEBA & AI Anomaly Detection Moderate Medium

Baseline normal user and entity behavior and flag the deviations automatically so analysts aren't hunting blind.

Automated Investigation & Response (SOAR) Moderate Medium

Automate the repetitive triage and response steps so a lean team can keep up with the alert volume.

API Security Monitoring Moderate Medium

See and stop attacks and abuse hitting my APIs — the blind spot my SIEM and WAF miss.

Dashboards, Visualization & Reporting Moderate Medium

Build dashboards and scheduled reports that make sense to analysts and to auditors alike.

Compliance & Audit Reporting Moderate Medium

Prove log retention and produce the PCI, HIPAA, and SOC 2 evidence auditors ask for without a fire drill.

Scale & Performance at High Volume Moderate Medium

Keep ingesting and searching as log volume grows without the platform falling over or costs exploding.

Ease of Deployment & Day-2 Operations Weak High

Stand it up and keep it running without a dedicated infrastructure team or deep Elasticsearch expertise.

Pick Your Spikes The audit tests all 12 capabilities, but competitive differentiation queries will emphasize 3. Which of these four Strong-rated capabilities best represents where Graylog wins deals?

  • Centralized Log Management & Collection
  • Fast Search & Incident Investigation
  • Log Parsing, Pipelines & Data Routing
  • Predictable, Volume-Based Pricing

→ Pressure-Test the Ratings Three checks: (1) Is the lone Weak accurate? We rated Ease of Deployment & Day-2 Operations weak from consistent reviewer complaints about setup and Elasticsearch expertise — does that still hold against "easy to run" competitors like Wazuh and Microsoft Sentinel, or has managed Graylog Cloud changed the story? (2) Is Threat Detection really only Moderate versus Splunk ES and Elastic Security, or are we underselling it? (3) Merge candidates — do UEBA & AI Anomaly Detection and Threat Detection & SIEM Correlation read as one capability to your buyers, or two? Merging them would consolidate their query coverage.

Pain Points

What Drives the Search

10 pain points — 5 high, 5 medium severity. The buyer language here is how we'll phrase queries, so the wording matters as much as the ranking.

Ingestion-based SIEM pricing makes budgets unpredictable High High

"Our Splunk bill is out of control — we're literally choosing which logs we can afford to keep"
Personas: CISO / VP Security · IT Ops Director · SOC Manager

Analysts buried in high-volume, low-signal alerts High High

"My analysts are drowning in alerts — we're going to miss the one that actually matters"
Personas: SOC Manager · Detection Engineer · CISO / VP Security

Lean teams can't operate a heavyweight SIEM High High

"We're a three-person security team — we can't babysit a SIEM that needs a full-time engineer just to keep it alive"
Personas: CISO / VP Security · SOC Manager

Slow log search delays incident containment High High

"When something's on fire I need answers in seconds, not to wait five minutes for a query to come back"
Personas: Detection Engineer · SOC Manager

Compliance evidence is a recurring manual scramble High High

"Every audit we burn a week pulling together log evidence to prove we're retaining and monitoring what we're supposed to"
Personas: Compliance / GRC Manager · CISO / VP Security

Deployment demands underestimated infra expertise Medium High

"It's powerful, but getting it configured and keeping it healthy takes more infra know-how than we bargained for"
Personas: IT Ops Director · Detection Engineer

Logs arrive in dozens of inconsistent formats Medium Medium

"Every source spits out logs differently and I spend half my time writing parsers just to make them searchable"
Personas: Detection Engineer · IT Ops Director

APIs are a security blind spot for SIEM and WAF Medium Medium

"We have no idea what's happening to our APIs until something breaks — our SIEM and WAF just don't see it"
Personas: CISO / VP Security · SOC Manager · Detection Engineer

Building detection content from scratch is slow Medium Medium

"I don't have time to hand-write detections for every MITRE technique — I need coverage that works on day one"
Personas: Detection Engineer · SOC Manager

Tool sprawl creates silos and duplicate cost Medium Medium

"I'm paying for three overlapping tools and still stitching data together by hand across IT and security"
Personas: CISO / VP Security · IT Ops Director

→ Rank and Refine Three checks: (1) Severity — five pains sit at high; is that the real order, or does one of the mediums (say, the API-security blind spot) actually cost you deals more than a listed high? (2) Buyer language — do these first-person lines sound like your buyers, or are we paraphrasing? The exact wording becomes the query text. (3) Inferred vs. real, and gaps — "tool sprawl" is inferred, not sourced from reviews; does consolidating three overlapping tools actually drive your deals? And are we missing pains we'd expect in log/SIEM deals — a painful rip-and-replace migration off an incumbent SIEM, storage/retention cost tradeoffs, or multi-cloud log ingestion gaps? Which of those show up in yours?

Site Findings

Technical Baseline (Layer 1)

The technical health of graylog.org as AI crawlers see it. These are engineering hand-offs — content recommendations come later in the full audit, once we know which gaps actually cost citations.

Start Here — Engineering One issue supersedes everything else in this document. graylog.org's WAF/bot-management layer returns HTTP 403 to non-browser clients site-wide and escalates to an IP-level 403 across all paths after roughly 11 requests — so GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, and Google-Extended are served a block instead of content. robots.txt and sitemap.xml also return 403, so crawl directives and the lastmod inventory are unreadable. Until engineering allow-lists the AI crawler user-agents and their IP ranges, no amount of on-page optimization can make Graylog visible in AI answers. This behavior appeared between the 2026-06-14 and 2026-07-23 crawls — engineering should also identify the WAF/CDN rule change that introduced it.

🔴 WAF/bot protection returns HTTP 403 to automated crawlers site-wide

What we found: graylog.org is fronted by aggressive WAF/bot-management (Cloudflare-style) that returns HTTP 403 to non-browser clients. robots.txt, sitemap.xml, and the HTML /sitemap/ page returned 403 to every request — from our fetch tool and from a browser-user-agent curl on the same host. Individual HTML pages were fetchable one at a time initially, but after ~11 requests the WAF escalated to an IP-level 403 across all paths: a fresh, never-requested product page and a browser-UA curl to /products/security/ both returned 403. This was not present in the prior crawl on 2026-06-14, when everything was freely retrievable — bot protection was tightened between then and now.

Why it matters: AI answer engines crawl with their own user-agents (GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Bytespider) from datacenter IP ranges. A WAF that 403s datacenter clients and rate-limits after a handful of requests serves these crawlers a 403 instead of content — making Graylog's pages functionally invisible to AI-powered search regardless of how good the on-page content is. This is the single highest-impact issue in the audit: it can zero out AI visibility upstream of every content optimization.

Business consequence: Queries like "best SIEM for lean security teams" or "predictable-pricing Splunk alternative" can only surface Graylog if AI crawlers can fetch graylog.org — while they receive a 403, competitors are cited in Graylog's place across every category and comparison query.

Recommended fix: Audit the WAF/CDN bot rules and explicitly allow-list the major AI crawler user-agents (GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Googlebot) and their published IP ranges so they receive 200s rather than 403/challenge. Confirm robots.txt and sitemap.xml are exempt from all bot challenges. Validate by curling each crawler's user-agent against the homepage, a product page, a feature page, and robots.txt/sitemap.xml, and by checking server logs for 403s served to those agents.

Impact: Critical Effort: 1–3 days Owner: Engineering Affected: Entire site

🟡 robots.txt and sitemap.xml return 403 — crawl directives and URL inventory unavailable

What we found: Both https://graylog.org/robots.txt and https://graylog.org/sitemap.xml (and the sitemap index and /sitemap/ HTML page) returned HTTP 403 to automated requests, including a browser-user-agent curl. The XML sitemap could not be retrieved, so the machine-readable URL inventory and per-URL lastmod timestamps — available in the prior crawl — are gone. Page discovery this run fell back to homepage/product-page navigation links plus site: web search.

Why it matters: Crawlers use sitemap.xml to discover pages efficiently and to read lastmod freshness signals; a 403 removes both, so deep pages (comparison and blog content most cited by LLMs) may go undiscovered and no page gets a recency signal from the sitemap. A robots.txt that returns 403 rather than 200 or 404 is ambiguous — some crawlers treat an unreachable robots.txt as a signal to back off or disallow, which can suppress crawling entirely.

Business consequence: The comparison pages most likely to be cited in "Graylog vs. Splunk" or "Graylog vs. Elastic" queries may never be discovered by AI crawlers if the sitemap that lists them keeps returning 403.

Recommended fix: Ensure /robots.txt and /sitemap.xml always return 200 to every user-agent (exempt them from WAF challenges and rate limits), keep the sitemap's <lastmod> values accurate, and declare the sitemap URL inside robots.txt. Re-test with multiple user-agents after the change.

Impact: High Effort: < 1 day Owner: Engineering Affected: robots.txt, sitemap.xml, child sitemaps

🟡 No recency signal on any page — undated content plus inaccessible sitemap

What we found: None of the 11 pages fetched live displays a visible published or updated date, and sitemap.xml (the lastmod source) is 403. Freshness is therefore undeterminable for every page. The content_marketing category — the 6 competitor comparison pages plus the "Mastering SIEM" guide — scores 0.20, and the category-weighted freshness average is 0.20 (red).

Why it matters: AI answer engines concentrate citations on recently-updated content: 76.4% of ChatGPT's most-cited pages were updated within the prior 30 days (ConvertMate, Q4 2025, ChatGPT-scoped), and AI-cited content runs 25.7% fresher on average than Google organic results (Ahrefs, August 2025). With no on-page dates and no readable sitemap lastmod, crawlers cannot establish recency for any Graylog page, so competitors' dated comparison content is preferred for citation. This regressed from the prior crawl because the sitemap is now WAF-blocked.

Business consequence: In the log management and SIEM category, undated comparison pages mean a query like "best Splunk alternative in 2026" is more likely to cite a competitor whose pages carry a visible recent date than to cite Graylog.

Recommended fix: Add visible "Last updated: <date>" stamps to all comparison and blog/guide pages (content_marketing), keep them current, and restore public access to sitemap.xml with accurate <lastmod> values. Prioritize the competitor comparison pages, which are the highest-leverage AI-citation surface.

Impact: High Effort: 1–2 weeks Owner: Content Affected: All comparison/blog pages; site-wide sitemap signal

🔵 Multiple H1s / non-semantic heading hierarchy on key landing pages

What we found: Several high-value pages render multiple top-level (H1) headings and stylistic rather than semantic nesting. The /why-graylog/ page returned seven H1-level headings ("Why Graylog? Because it works.", "The World Today", "What Our Customers Say", "One Platform for Full Visibility", "The Graylog Advantage", etc.), and the /use-cases/sec-ops/ template mixes four H1-level headings with the section H2s. The homepage and /open-vs-paid/ show similar section-as-heading patterns.

Why it matters: Multiple H1s and stylistic heading use blur the document outline that LLMs rely on to segment a page into self-contained, citable passages. When every section is an H1, no single heading carries the page's primary claim, weakening passage extractability on exactly the executive-facing pages meant to win the CISO buyer.

Business consequence: On the executive-facing /why-graylog/ page meant to win the CISO, a blurred heading outline makes it harder for an engine to extract a clean, citable answer to "why Graylog," handing cleaner-structured competitor pages the citation.

Recommended fix: Enforce a single descriptive H1 per page and a logical H2→H3 outline on the affected templates (why-graylog, use-case, and homepage templates). Convert stylistic H1s to H2/H3 with noun-phrase labels that state the section's claim.

Impact: Medium Effort: 1–3 days Owner: Engineering Affected: /why-graylog/, /use-cases/sec-ops/, homepage, /open-vs-paid/

Manual Verification Checklist

The following items could not be assessed through our analysis method (rendered markdown) — several were also cut short when the WAF blocked further requests mid-analysis. We recommend your engineering team verify these manually before the validation call.

Client-side-rendering status not fully verifiable

What to check: Every page we did retrieve returned full server-rendered body content (no blank-shell/CSR symptom), but because most pages became unreachable (403) mid-analysis, client-side-rendering behavior could not be confirmed site-wide. Our fetch tool also returns rendered markdown, so framework/CSR signals in the raw HTML aren't directly visible.

Recommended action: Verify rendering with JavaScript disabled and with a crawler that fetches raw HTML (Screaming Frog in list mode, or "view-source" / Google's URL Inspection "rendered vs. raw HTML") across a sample of product, feature, comparison, and use-case templates.

Effort: < 1 dayOwner: Engineering

Structured data (JSON-LD schema) not assessable from rendered output

What to check: Our analysis method returns rendered markdown, not raw HTML, so JSON-LD/schema.org markup is not visible — schema_coverage is null for all 36 inventoried pages. Verify rather than assume presence or absence.

Recommended action: Verify with a structured-data testing tool (Google Rich Results Test / Schema Markup Validator) on a sample per template. Add the most specific applicable type — Product on /products/*, FAQPage where FAQs exist, Article with datePublished/dateModified on comparisons and blog posts (which also helps the freshness gap).

Effort: 1–3 daysOwner: Engineering

Meta descriptions and Open Graph tags not assessable from rendered output

What to check: Meta description and Open Graph/Twitter card tags live in the raw HTML <head> and aren't visible in rendered markdown, so meta_description is null and has_og_tags is unverified across the inventory. This is a tooling limitation, not a confirmed absence.

Recommended action: Verify with a social-preview/inspection tool or view-source on a sample. Ensure each commercial page has a unique, benefit-specific meta description and complete OG tags (title, description, image, url).

Effort: < 1 dayOwner: Marketing

Site Analysis Summary

Pages analyzed 36 inventoried (only ~11 fetched live before WAF block)
Commercially relevant pages 35
Heading hierarchy (avg) 0.68
Content depth (avg) 0.62
Passage extractability (avg) 0.65
Freshness (weighted) 0.20 · content marketing 0.20 · product: unable to assess (28 undated) · structural: unable to assess (1)
Schema coverage (avg) Unable to assess (36 pages unscored)

Partial Sample The WAF cut this analysis short: of 36 inventoried pages, only ~11 were fetched live before the IP-level block, 29 of 36 pages could not be scored for freshness, and all 36 are unscored for schema. Treat these averages as a partial read — once crawler access is restored, a clean re-crawl will give fuller coverage and firmer scores.

Next Steps

What Happens Next

Why Now The window to establish AI visibility in the SIEM and log-management category is open, not permanent:

  • AI search adoption is accelerating — 87% of B2B software buyers say AI chatbots are changing how they research vendors, and half now start research in a chatbot rather than Google (G2, October 2025).
  • Early citations compound: domains that AI platforms learn to trust now get cited more frequently as training and retrieval data accumulate.
  • Competitors who establish GEO visibility first create a structural disadvantage for late movers — being the cited answer is self-reinforcing.
  • Log management and SIEM is still early-innings in GEO optimization — acting now means competing against inaction, not against entrenched strategies.

The full audit will measure Graylog's citation visibility across real buyer queries in the log management and SIEM space — from "best Splunk alternative with predictable pricing" and "SIEM a lean security team can actually run" to "Graylog vs. Elastic for detection engineering." You'll see exactly which queries return answers that include your competitors but not Graylog — and what it would take to appear in them. Fixing the crawler-access and freshness issues flagged above first means the audit measures an accessible site rather than a blocked one, so the results reflect your content, not your WAF.

01

Validation Call

45–60 minutes. We walk through this document together and lock the inputs — confirm personas, competitor tiers, feature strengths, and pain point priorities.

02

Query Generation & Execution

We build the buyer query set from the validated inputs and run it across the selected AI platforms, capturing who gets cited for each query.

03

Full Audit Delivery

Visibility analysis, competitive positioning, and a three-layer action plan — prioritized by which gaps actually cost you citations.

Start Now — Engineering Three technical fixes your engineering team can start before the call, independent of the audit results: (1) Allow-list the major AI crawler user-agents and IP ranges in the WAF/CDN so they get 200s, not 403s — and confirm robots.txt and sitemap.xml are exempt from all bot challenges (both currently return 403, so verify and resolve that first). (2) Restore a public, accurate sitemap.xml with lastmod values so crawlers can discover deep pages and read recency. (3) Enforce a single semantic H1 and a logical H2→H3 outline on the why-graylog, use-case, and homepage templates. These don't depend on the rest of the audit and will improve your baseline visibility before we even measure it.

Before the Call

Your Pre-Call Checklist

Two jobs before we meet. The questions on the left require your judgment — no one knows your business better than you. The engineering tasks on the right don't require the call at all.

Questions for You
Does a dedicated compliance / GRC buyer (Susan Albright) actually sit in Graylog's buying committee?
If wrong: we drop the persona and its entire PCI/HIPAA/SOC 2 compliance-evidence query cluster.
Is free, open-source Wazuh a real head-to-head primary, or a DIY / build-vs-buy alternative?
If secondary: ~5 head-to-head queries shift out of the direct-comparison set.
Is API Security a separate buying motion, or part of the core SIEM evaluation?
If separate: it becomes its own query cluster instead of being tested inside SIEM queries.
Is Ease of Deployment still a genuine Weak spot, or has managed Graylog Cloud changed it?
If outdated: we stop conceding "easy to run" ground to Wazuh and Microsoft Sentinel in queries.
Is "tool sprawl" a real deal driver, and are we missing pains like SIEM migration or retention cost?
If wrong: we re-weight the pain set that determines how queries are phrased.
Which 3 of the 4 Strong capabilities best represent where Graylog wins deals?
If unspecified: differentiation queries emphasize the wrong spikes.
Does SOC Manager Elena Vasquez hold practical veto in a lean team?
If yes: we reclassify her decision-maker and add operability queries to the decision set.
Does IT Ops Director David Okonkwo control the infrastructure/logging budget line?
If yes: he moves to decision-maker and cost-predictability queries gain a second buyer.
Does Detection Engineer Priya Raman initiate evaluations bottom-up via Graylog Open?
If yes: we add practitioner-stage queries (detection content, parsing, search speed).
In a lean org, does CISO Marcus Bell run the technical evaluation himself?
If yes: executive and practitioner queries collapse into one cluster.
Are we missing a vendor (QRadar, Cribl, CrowdStrike Falcon LogScale) or listing noise (Grafana Loki, Log360)?
If wrong: the head-to-head competitor set — and its query budget — is misallocated.
Which roles beyond the five listed show up in your deals (MSP/MSSP, DevOps lead, Procurement)?
If any do: they warrant their own dedicated query cluster.
For Engineering — Start Now
Allow-list AI crawler user-agents and IP ranges in the WAF/CDN so they receive 200s, not 403s.
The single highest-impact fix — it can zero out AI visibility upstream of everything else.
Restore robots.txt and sitemap.xml to a 200 for every user-agent, with accurate lastmod values.
Both currently 403 — restores crawl directives, page discovery, and recency signals.
Enforce a single semantic H1 and H2→H3 outline on why-graylog, use-case, and homepage templates.
Sharpens passage extractability on the executive pages meant to win the CISO.
Verify client-side-rendering behavior with JavaScript disabled across a sample of templates.
Confirms crawlers that don't run JS still see full content.
Verify JSON-LD/schema markup with a structured-data testing tool on a sample per template.
Confirms whether Product/FAQPage/Article schema is present or missing.
Alignment

We're Aligned On

This isn't a contract — it's a shared understanding. The audit runs against what's below. If something changes between now and the call, we adjust. The goal is to make sure we're asking the right questions for the right buyers against the right competitors.
Already Confirmed
Competitive set — 5 primary (Splunk, Elastic Security, Wazuh, Microsoft Sentinel, Sumo Logic) + 5 secondary (Datadog, Exabeam/LogRhythm, Devo, ManageEngine Log360, Grafana Loki)
Persona set — 5 personas: 1 decision-maker, 1 evaluator, 3 influencers
Feature taxonomy — 12 capabilities (4 strong, 7 moderate, 1 weak) with outside-in strength ratings
Pain point set — 10 buyer frustrations (5 high, 5 medium) with severity ratings
Layer 1 technical audit — 7 findings logged (1 critical, 2 high, 4 medium/low), engineering notified
Decided at the Call
Crawler access — confirm the WAF now serves AI crawlers 200s (not 403s) before queries run; until then the audit would measure a blocked site
Compliance / GRC persona (Susan Albright) — is a dedicated compliance buyer really in the committee, or is this a security + IT purchase? (inferred persona; a full query cluster rides on it)
Wazuh tier — primary head-to-head vs. secondary DIY alternative (shifts ~5 head-to-head queries)
Feature overweighting — which 3 of the 4 Strong capabilities to emphasize (our default: Predictable Pricing, Fast Search, Centralized Log Management — each maps to a high-severity pain)
Pain point prioritization — top 3 buyer problems to test first (cost explosion, alert fatigue, lean-team capacity lead on severity × persona breadth)
Persona corrections (Elena's veto, David's budget control) and competitor set adjustments (missing vendors / listing noise)
Client
Date