Operational visibility
Agent observability starts after the model response
Traditional LLM telemetry often focuses on prompts, responses, latency and tokens. Those signals matter, but an agent can continue acting after a response is generated. It may choose a tool, query data, edit a file, run a command, contact another agent or wait for human approval. Observability must follow that execution path if the team wants to understand the real outcome.
Agnys captures model and operational events in the same session trace. Teams can watch activity as it happens, search historical fields and replay a run without switching among vendor consoles. The shared record connects runtime operations, security investigation, cost review and governance evidence rather than producing a separate dashboard for each concern.
- Model calls and usage context
- Tool, MCP and API activity
- File edits and shell commands
- Human approvals and session outcomes
Live action feed
Watch agent actions in the order they occur
A live feed makes an agent’s work visible at the event level. Instead of learning about a problem from a customer or a downstream system, operators can see the calls and side effects that make up the run. Each event retains the agent, timestamp, type and trace relationship needed to move from a signal to the surrounding context.
This is useful during rollout, when teams are still learning how a production agent behaves outside a controlled demonstration. It also supports day-to-day operations: reviewing expensive sessions, checking whether a tool integration is producing unexpected retries, and confirming that an agent stayed inside its intended workflow.
- Filter activity by agent, session, event type and time range
- Move from a notable event directly into its full session
- Keep model and side-effect activity on the same timeline

Behavioral DNA
Detect changes relative to each agent’s own baseline
A fixed threshold can catch obvious failures, but many agent incidents are contextual. Ten tool calls may be normal for one agent and highly unusual for another. Agnys learns patterns such as tool mix, event sequence, timing, cost and approval behavior for each agent, then scores new activity against that baseline.
Behavioral drift is a prompt to investigate, not a declaration that the agent is malicious or broken. A product release, new customer workflow or legitimate data migration can change normal behavior. Agnys keeps the evidence around the deviation so a qualified operator can decide whether to accept the new pattern, refine the baseline or escalate the event.
- Unusual tool selection or event sequence
- Changes in volume, timing or cost patterns
- Shifts in oversight and approval behavior
- Scope expansion into files or systems not normally touched
Anomaly response
Give alerts enough evidence to be actionable
An alert without context creates another investigation queue. Agnys detections point back to the event and session that produced the signal. Operators can inspect the preceding model call, the tool arguments, related approvals and the resulting side effect before choosing a response.
Rules can target known concerns while anomaly scoring highlights deviations that a static rule may not anticipate. Detections can be tagged to relevant threat frameworks and sent through configured channels such as email, Slack or PagerDuty. The preview and live evaluation use the same underlying event fields so teams can test a rule before relying on it operationally.
- Investigate prompt injection signals and unexpected scope expansion
- Identify destructive command bursts or repeated tool failures
- Route alerts with the trace context needed for triage
Search and replay
Ask precise questions about historical behavior
Operations teams need more than a chronological feed. Agnys Search Lab provides a query language across captured fields, with pipeline statistics, sequence operators, saved searches and alert rules. A team can look for a particular tool, compare activity by agent, find sessions above a cost threshold or identify an event sequence that preceded earlier failures.
Once a result is found, forensic replay rebuilds the session from trigger to outcome. That combination supports incident response and product improvement: search across the fleet to find a pattern, then study individual sessions to understand why it occurred. Saved searches make the investigation repeatable as the agent or deployment changes.
- Query model, tool, file, command, approval and cost fields
- Use sequence-aware searches to find multi-step behavior
- Turn a useful investigation into a saved search or alert
Cost context
Tie spend to the behavior that created it
Token totals show how much an agent consumed, but not whether the spend produced useful work. Agnys computes model cost at ingest and keeps it alongside the session events. A cost spike can therefore be examined with the calls, tools, retries and outcome that generated it.
This helps teams separate legitimate high-value sessions from loops, excessive context and failing integrations. Cost becomes another operational signal rather than a monthly surprise. Rollups support dashboards and retention planning, while the trace remains available when a team needs to explain a specific outlier.
Cross-team use
A shared operational picture for engineering and governance
AI builders need debugging detail. Security teams need attribution and unusual behavior. Compliance leaders need reviewable records and control evidence. Business owners need confidence that agents are operating inside the workflow they approved. These groups often work from different tools and reach different conclusions about the same run.
Agnys uses one tamper-evident event record across live monitoring, replay, anomaly detection, search, cost tracking and evidence export. Each team can use the view appropriate to its role while referring back to the same captured events. That reduces the gap between an engineering explanation and the evidence presented during customer assurance or an audit.