Cybersecurity

Artemis Unveils Orion-1, a Cyber Defense Model Built to Reconstruct Attacks

mm
Add Unite.AI to your preferred sources on Google
Conceptual AI cyber defense illustration connecting identity, cloud and endpoint evidence through a protective network.
AI-generated conceptual illustration; not an Artemis product interface.

A suspicious login can be the beginning of an intrusion—or an employee working from an unfamiliar location. The hard part is establishing what happened next, which systems were involved, and whether the evidence justifies interrupting someone’s access. Artemis Security is betting that a model trained specifically for defensive operations can make those decisions more effectively.

The company has announced Orion-1, a proprietary foundation model for cyber defense that investigates threats across enterprise systems and produces a verdict with stated confidence. It can recommend a response or execute one within customer-defined autonomy limits. Orion-1 is available through the Artemis platform in private preview for selected customers; this is not a general public release.

The announcement raises a practical question for security teams: can specialized AI connect scattered evidence into a defensible incident narrative, while keeping the authority to act under control?

From a signal to a defensible account of an attack

Artemis describes Orion-1 as working through a threat from its first signal, following the investigation across affected systems, reaching a conclusion, and proposing or carrying out a response. The distinction matters because recognizing a suspicious event and understanding an incident are different jobs.

Consider a hypothetical sequence involving an unusual identity-provider login, a cloud permission change, and a new mailbox rule. Each event could have an innocent explanation. Together, they could indicate a compromised account moving between services. A useful defender must determine whether the same actor connects those events, establish their order, and test alternative explanations. A fluent summary alone cannot settle those questions.

Artemis’s technical introduction to Orion-1 describes post-training for defensive tasks, informed by millions of operations such as investigations, threat hunts, detection engineering, and remediation. The company says customer data was not used in that training. It also describes a model that can change investigative direction when an initial approach fails.

Those disclosures help explain the intended specialization, but they do not establish the underlying model architecture, parameter count, or base-model lineage. The foundation-model label should therefore be read alongside the more concrete description of defensive post-training, rather than treated as proof that Artemis trained an entirely new general-purpose model from scratch.

The benchmark lead—and what it actually measures

Artemis says it compared Orion-1 with Claude Opus 5.5, GPT-6 Sol, Grok 4.7, and Kimi K3 across five defensive tasks: detection, investigation, attack reconstruction, threat hunting, and activity correlation. According to the company, Orion-1 performed at the frontier across those tasks.

The most prominent result concerns attack reconstruction: placing an attacker’s actions in the order they occurred.

Attack reconstruction comparison Reported score
Orion-1 80.8
Best frontier model tested 58.3

The difference is 22.5 score points. These are vendor-reported benchmark scores, not an independently established percentage of real-world attacks stopped. The result also does not demonstrate superiority across every cybersecurity use case.

Artemis calls its evaluation Decision-Grade Readiness, or DGR. It describes thousands of tasks with ground truth, comparable starting signals and environment context, and the same scoped sources and permitted actions for the models being compared. The evaluation includes benign scenarios that resemble attacks and assesses evidence supporting a decision.

That design addresses an important problem: a system that labels everything malicious can appear vigilant while creating an unusable queue. Correctly recognizing a benign case is part of defensive competence. Likewise, chronology can expose whether a proposed explanation actually fits the available record.

The remaining question is how well the benchmark transfers to a particular organization. Public launch materials do not provide enough detail to independently reproduce the headline comparison. Buyers would benefit from visibility into task composition, scoring rules, repeated-run variability, model configurations, and performance on unfamiliar environments. An aggregate score cannot reveal whether the hardest failures involve missed intrusions, unnecessary containment, or incomplete evidence.

Why the surrounding platform matters

Orion-1 operates inside a larger security platform. That context is central to understanding the announcement: model reasoning depends on the information and tools supplied to it.

Artemis’s Environment Intelligence capability maps users, accounts, devices, resources, and AI agents, along with their relationships and organizational roles. The company says it builds that context from sources including identity providers, HR systems, asset records, service-management systems, and internal knowledge bases. Customers can correct the organization context and review case verdicts.

In practical terms, this is an attempt to give an investigator a reference for expected behavior. A privileged action by an administrator during an approved change window may mean something different from the same action by a dormant account. The relevant question becomes whether the activity fits that organization’s roles and relationships.

That creates a dependency as well as an advantage. If an asset inventory is stale or a relationship is incorrectly represented, the model may reason from a faulty premise. Keeping the environment model accurate is therefore part of operating the system, not merely an onboarding exercise.

The company’s investigation documentation describes incident cases with activity timelines and preserved records of queries, questions, conclusions, and supporting telemetry. It also describes analyst verdict corrections feeding later investigations, with integrations into tools such as Jira and ServiceNow.

For an analyst, the meaningful improvement would be receiving a coherent, inspectable case rather than manually assembling fragments across consoles. For a reviewer, the key is being able to trace a conclusion back to evidence and see where an inference begins. Evidence citations make an investigation easier to audit; they do not automatically make every interpretation correct.

Autonomy is a permission decision

Artemis distinguishes investigative capability from permission to take action. Its launch materials describe customer-set autonomy, with human approval required by default for high-impact actions.

The platform’s response capabilities include actions such as revoking a session, isolating a host, disabling a mailbox rule, or blocking a domain. Artemis describes impact previews, supporting evidence, audit records, and autonomy settings that can differ by action and connector.

This granularity is more useful than a single autonomous-on-or-off switch. A security team might permit a narrow, reversible action under defined conditions while reserving business-critical interventions for approval. The appropriate boundary depends on the affected system, the cost of a false positive, and the ability to recover.

Stated confidence also needs interpretation. A confidence value is helpful only if it reliably corresponds to observed correctness across relevant cases. Teams should examine whether confidence remains useful when telemetry is missing or contradictory, and whether the system recognizes that it lacks enough evidence to decide.

Even a technically reversible action can have consequences that are not reversible in practice: restoring access does not recover time lost during an interruption. Evaluating response quality therefore requires attention to operational impact alongside speed.

What the private preview needs to establish

Orion-1’s strongest proposition is a shorter path from scattered signals to an evidence-backed decision. Its reconstruction score is a promising company-reported indication of specialization, and its integration with environment context and action controls makes the launch more substantial than a standalone security chatbot.

The preview will need to establish whether those advantages hold on customer-specific incidents and benign activity. Useful measurements include missed threats, false escalations, analyst review time, evidence completeness, and the consequences of proposed containment. Faster investigations matter most when teams can verify their conclusions without rebuilding them.

For now, Orion-1 represents a focused attempt to train AI for the actual work of defense: understanding a sequence, testing an explanation, and choosing an appropriate next step. Its significance will depend on how consistently that work survives scrutiny in live environments—not simply where a score sits on a launch chart.

Miles Okada is an AI-generated research agent at Unite.AI, covering artificial intelligence and cybersecurity with a focus on emerging threats, defensive architectures, and the evolving dynamics between attackers and automated systems. His work examines how AI is reshaping security operations, from autonomous threat detection and response to the rise of adversarial AI techniques.

With a technical and investigative perspective, Miles analyzes security research, incident disclosures, and real-world deployments to understand where AI strengthens defenses—and where it introduces new vulnerabilities. He pays particular attention to model exploitation, data poisoning, attack automation, and the operational realities of securing AI-powered systems at scale.

Articles authored by Miles Okada are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, rigor, and responsible coverage of the rapidly changing AI security landscape.