The evolution of AI at Elastic Security — three years from a chat sidekick to agents that investigate, reason, and act.
James Spiteri · Elastic Security
presenter: click a card per raised hand · shift-click to undo
ChatGPT moment. Natural language becomes the interface — AI Assistant answers questions, explains alerts, writes queries. Human does everything else.
LLMs aimed at specific painful jobs: Attack Discovery triages alerts in bulk; Automatic Import, Migration & Troubleshooting kill weeks of toil. Output lands in real UIs, not chat.
Agents with tools, reasoning you can read, workflows, MCP & A2A. From answering questions to running investigations end-to-end — with the analyst in control.
ChatGPT had just taken the world by storm — natural language democratized the technology. We shipped the Elastic AI Assistant to find out what value it could actually deliver in a SOC.

The actual 2023 UI — my own connector, 32K tokens and all. Promising, but a long way from carrying real SOC weight.
Context windows → tool use → readable reasoning: each capability unlocked the next product era.

Bulk alert triage → discoveries with attack chains, entities, and MITRE ATT&CK mapping.
LLM-powered triage over alerts in bulk, with findings in a dedicated interface — not a conversation. Claude 3 had just landed, and the results told us we could keep pushing.
Chat is an interface, not the product. Put the model inside the workflow and give the output a real home in the UI.
Alert fatigue is a volume problem. Machines read 100 alerts the way analysts read one — connecting them into attack chains automatically.
Attack Discovery's results made one thing obvious: quality drifts. New models, new prompts, new data shapes — without evaluations, you're shipping vibes.

Real eval heatmap: adjacent prompt variants swinging from ~95% to ~25% accuracy.
Point it at any log source; it builds the custom integration and pipeline. Weeks of onboarding → minutes.
Translates rules & dashboards from your legacy SIEM into Elastic equivalents — the #1 migration blocker, automated.
Diagnoses Elastic Defend deployment issues automatically instead of hours of doc-spelunking.
Assistant re-architected on LangChain LangGraph — first agent graph with knowledge sources, sub-agents for ES|QL query generation. Attack Discovery pivoted to agents too.

Global Generative AI Infrastructure & Data Partner of the Year — announced at re:Invent 2024. Among the first 15 partners awarded the AWS GenAI Competency.
Named #2 of the "Top 5 LangGraph Agents in Production 2024" — recognized as one of the first companies anywhere to ship a production AI agent.
Five-year strategic collaboration agreement with AWS (2025) — accelerating agentic AI, MCP, and agent-to-agent interoperability. 2025 GenAI Partner award global finalist.

Keep joined forces with Elastic — automation for agents, agents for automation.
Rebuilt with our search & platform teams: agents defined once in the stack — not bolted onto one app.
Deterministic YAML workflows become tools agents can call — LLM judgment where you want it, rules-based certainty where you need it. GA in Serverless & 9.3.
Agents that pull context, run queries, and build the picture across your data — reasoning shown, analyst in charge.
Entity risk explained in plain language — why this host, why this score, what to do next — right in Entity Analytics.
Agent Builder + Workflows + MCP/A2A across the platform: compose the SOC automation you actually want.

9.3: AI-generated entity summaries with recommended actions, inside the analyst workflow.

The full loop: Attack Discovery finds it → workflow opens a case → agent investigates → findings + reasoning land in Slack.
Native automation in Elastic Security — triage, enrichment, response, case management. No standalone SOAR required.
Five SOC skills: alert triage, detection rule authoring, entity investigation, threat hunting, anomaly analysis — chained in one investigation. Agent Builder GA in Security.
Precision Entity Identification, Entity Resolution, Dynamic Watchlists, hunting leads — one authoritative record per person, so agents reason over accurate context.
9.5 is around the corner — more skills incoming: detection emulation, binary analysis, alert dedup.
Identical-looking prompts swing 95% → 25%. Measure or you're shipping vibes.
Aim models at specific, hated jobs and give outputs a real home in the UI.
Scoped tools, workflow guardrails, explicit boundaries. Agents earn write access.
Analysts trust what they can audit. Show the thought process, always.
500+ conversations, one signal: customization and insight beat autonomy theater.
James Spiteri · Elastic Security — thank you. Questions: find me after 🤝