LLMOps and Observability: Tracing, Metrics, Drift and Feedback Loops

By Aakash Ahuja··33 min read

LLMOps and observability exist to answer one question about a production LLM or agent system: what exactly ran, what happened inside it, whether the result was good, what it cost, why it failed, and how to change or roll it back safely. Traditional observability tells you whether a system is healthy; LLM observability must also tell you whether it is behaving correctly. This article covers versioning of every behaviour-changing artifact, OpenTelemetry tracing with the GenAI semantic conventions, the metrics hierarchy, drift detection, incident response, rollback and the feedback loop that turns production into evaluation data.

The whole topic reduces to one question:

Once an LLM or agent system is in production, how do I know exactly what version ran, what happened inside it, whether it produced a good result, how much it cost, why it failed, and how I can safely change or roll it back?

Traditional application observability answers “Is the system healthy?” LLM observability must additionally answer “Is the system behaving correctly?”

That distinction is the centre of the architect's view of this domain.


14.1 What is LLMOps? The mental model

LLMOps is the discipline of versioning, tracing, measuring, evaluating and safely changing LLM and agent systems once they are in production.

Think of production GenAI operations as six connected capabilities:

CapabilityQuestion answered
VersioningWhat exactly was running?
TracingWhat exactly happened?
MetricsHow well did it perform?
EvaluationWas the output actually good?
Monitoring / driftIs behaviour changing over time?
Incident / rollbackHow do we recover when something goes wrong?
A mature AI platform therefore needs something like:

                       ┌────────────────────┐
                       │   User / Channel   │
                       └─────────┬──────────┘
                                 │
                                 ▼
                        ┌──────────────────┐
                        │ API / AI Gateway │
                        └────────┬─────────┘
                                 │
                     trace_id / tenant / app
                                 │
                                 ▼
                    ┌────────────────────────┐
                    │ Agent / RAG Runtime    │
                    │                        │
                    │ Agent Run              │
                    │ ├── Policy check       │
                    │ ├── Retrieval          │
                    │ │   └── Reranker       │
                    │ ├── LLM call           │
                    │ ├── Tool call          │
                    │ ├── Verification       │
                    │ └── Human approval     │
                    └───────────┬────────────┘
                                │
                  OpenTelemetry instrumentation
                                │
                                ▼
                   ┌─────────────────────────┐
                   │ OTel Collector / Bus    │
                   └───────────┬─────────────┘
                               │
              ┌────────────────┼─────────────────┐
              ▼                ▼                 ▼
          Traces           Metrics             Logs
              │                │                 │
              └────────────────┼─────────────────┘
                               ▼
                    Observability Platform
                               │
         ┌─────────────────────┼────────────────────┐
         ▼                     ▼                    ▼
     Dashboards             Alerts             Evaluation
     Latency                SLOs               Quality
     Tokens                 Errors             Safety
     Cost                   Drift              Grounding

OpenTelemetry is particularly important because its semantic conventions create common names and meanings for telemetry across frameworks and platforms. OpenTelemetry's current semantic-conventions documentation is version 1.44.0, and the GenAI conventions are now maintained in a dedicated GenAI semantic-conventions repository. (OpenTelemetry)


14.2 What is the LLM application lifecycle?

The LLM application lifecycle is the path from design through evaluation, staged rollout and production feedback, applied not to one artifact but to many that change independently.

Traditional DevOps is roughly:

code → build → test → deploy → monitor

LLMOps introduces several additional independently changing artifacts:

Code
Prompt
Model
Embedding model
Dataset
Knowledge corpus
Retrieval configuration
Reranker
Tool schemas
Agent graph
Policies
Evaluation suite
Infrastructure

So the lifecycle becomes:

Design
   ↓
Develop
   ↓
Prompt / Agent experimentation
   ↓
Offline evaluation
   ↓
Security / safety evaluation
   ↓
Integration testing
   ↓
Staging
   ↓
Shadow / canary
   ↓
Production
   ↓
Online evaluation
   ↓
Production feedback
   ↓
Drift / incident detection
   ↓
Improve → repeat

The architecturally important concept is lineage.

For every production response I ideally want to reconstruct:

response_id
   ↓
application_version = 4.8.1
agent_version       = 12
prompt_version      = 37
model               = model-X/version-Y
embedding_version   = 7
retrieval_config    = 18
dataset_version     = 46
policy_version      = 11
evaluation_version  = 23
deployment_version  = prod-20260818.4

That gives you reproducibility and RCA.


14.3 How should you version prompts?

Prompt versioning means storing every prompt as a registered, versioned artifact with an owner, evaluation evidence and a rollback target, rather than as a string inside application code.

Treat prompts as production artifacts, not strings buried in application code.

A prompt registry should track:

FieldExample
Prompt IDinvoice_extraction
Versionv17
TemplateSystem + task template
Parameterstemperature, output schema
OwnerFinance AI team
CommitGit SHA
Created byuser/service identity
Created attimestamp
Evaluationeval-run-2187
Statusexperimental/candidate/production
Model compatibilityapproved models
Rollback versionv16
Why?

Because this:

"Summarise the contract."

becoming:

"Summarise the following contract, focusing on obligations,
termination and liability..."

looks like a tiny configuration change but can fundamentally change system behaviour.

Modern LLM lifecycle systems therefore associate prompt versions with evaluations and traces; for example, MLflow's Prompt Registry supports evaluating prompt versions across models and datasets while maintaining experiment history. (MLflow AI Platform)

Enterprise pattern

Do not deploy:

prompt = latest

Prefer:

application
   ↓
prompt alias: production
   ↓
prompt version: 17

Then promotion is controlled:

v18
 ↓
offline eval
 ↓
security eval
 ↓
shadow
 ↓
canary
 ↓
production alias → v18

Rollback is simply:

production alias → v17

14.4 Model versioning

Model versioning must record more than the marketing model name.

Track:

provider
model family
exact model/deployment identifier
region
endpoint
model revision where available
context window
parameters/configuration
temperature
top_p
max tokens
reasoning settings
structured-output mode
fallback model
routing policy

Why?

Because:

same prompt
+ different model
= potentially different application

A provider can also expose models through intermediary services or proxies, so identifying both requested and actual model/provider matters. OpenTelemetry's GenAI attributes distinguish requested/response model information and provider information for this reason. (OpenTelemetry) Fallback models and routing policies are where this gets complicated in practice; the model gateway and routing layer is where that identity should be resolved and recorded.


14.5 Dataset versioning

Three datasets often exist separately:

Training / fine-tuning dataset
Evaluation dataset
Production RAG knowledge corpus

All need lineage.

For RAG, for example:

Corpus v41
 ├── source documents
 ├── parser version
 ├── OCR version
 ├── chunker version
 ├── embedding model
 ├── metadata schema
 └── index version

If retrieval suddenly deteriorates, you need to distinguish:

Did documents change?
Did parsing change?
Did chunking change?
Did embeddings change?
Did indexing change?
Did the query distribution change?

Without dataset/index lineage, those become extremely difficult to separate.


14.6 Evaluation versioning

An evaluation score is meaningless without knowing how it was calculated.

Track:

evaluation_suite_id
dataset_version
evaluator_version
judge_model
judge_prompt
metric definitions
thresholds
sampling strategy
test configuration
timestamp

Suppose:

Groundedness = 91%

Then someone changes the judge prompt.

Next week:

Groundedness = 96%

That does not necessarily mean the application improved.

The measurement system changed.

Therefore:

Application versioning
AND
Evaluation versioning

are equally important. How the evaluation suites themselves are built is covered in Evaluating LLM, RAG and Agent Systems.


14.7 Deployment versioning

A deployment should identify the complete AI configuration bundle, not just Docker image version.

Think:

AI Deployment Manifest

application: claims-agent:4.12
agent_graph: v27
prompt: v16
primary_model: model-A
fallback_model: model-B
embedding_model: embedding-v4
retriever: config-v13
reranker: reranker-v2
policy_pack: v8
tool_registry: v21
evaluation_gate: eval-v19

This creates an immutable deployable unit.


14.8 Feature flags

Feature flags become extremely useful for GenAI because many changes should not require redeployment.

Example:

use_new_reranker = true
enable_memory = false
new_prompt_v18 = 10% traffic
enable_agent_verifier = tenant_A_only
use_model_B = geography_EU

They support:

A/B tests
canaries
tenant-specific rollout
model migrations
emergency kill switches
cost experiments
new tool exposure

But be careful:

Feature flags themselves must become part of the trace.

Otherwise the system says:

version = 4.2

but two users received completely different behaviour because their flags differed.


14.9 Experiment tracking

Experiment tracking answers:

What configuration generated these results?

An experiment should link:

prompt
model
model parameters
retriever configuration
dataset
evaluation suite
metrics
results
cost
latency
developer
commit

Then you can compare:

ExperimentQualityP95 latencyCost/request
GPT-A + prompt 1492%3.4s$0.024
GPT-B + prompt 1491%1.8s$0.008
GPT-B + prompt 1693%1.9s$0.009
An architect should immediately see:

Highest-quality model does not necessarily equal best production configuration.

The real optimization is often:

maximize

quality × reliability

subject to

latency ≤ SLO
cost ≤ budget
security ≤ risk tolerance

14.10 What is traceability in an AI system?

Traceability means being able to reconstruct a single AI request end-to-end.

For example:

User: "Can supplier X be approved?"

trace_id = abc123

14:32:01 request received
14:32:01 policy check passed
14:32:02 supplier profile retrieved
14:32:02 sanctions DB queried
14:32:03 6 documents retrieved
14:32:03 reranker selected 3
14:32:04 LLM reasoning invoked
14:32:07 tool invoked: supplierRiskScore
14:32:08 LLM produced proposed decision
14:32:09 verifier detected missing tax certificate
14:32:10 human approval requested
14:33:41 human rejected

You should then be able to ask:

Why did the agent propose approval?
Which documents did it see?
Which model produced the result?
What prompt version?
What tools ran?
What permissions did they use?
How many tokens?
What did it cost?
How long did each step take?
Which policy version applied?

That is production-grade AI traceability.


14.11 Distributed tracing

Distributed tracing carries one correlation context across every service a request touches, so the hops can be stitched back into a single transaction.

Enterprise AI rarely lives in one service.

A request might cross:

Web application
     ↓
API Gateway
     ↓
Agent Service
     ↓
Retrieval Service
     ↓
Vector DB
     ↓
Reranker
     ↓
Model Gateway
     ↓
LLM Provider
     ↓
Tool Gateway
     ↓
SAP
     ↓
Approval Service

Every hop must preserve correlation:

trace_id
span_id
parent_span_id

Without this you end up with twenty sets of logs and no reliable way to reconstruct the transaction.


14.12 What is OpenTelemetry and why use it for AI?

This is the standard to build on.

OpenTelemetry is a vendor-neutral observability instrumentation standard/ecosystem for collecting and exporting telemetry such as:

Traces
Metrics
Logs

Its semantic conventions standardize names and meanings for attributes, spans, metrics and related telemetry, making cross-platform correlation much easier. (OpenTelemetry)

Typical architecture:

Application
    │
OpenTelemetry SDK
    │
    ├─ traces
    ├─ metrics
    └─ logs
    │
    ▼
OTel Collector
    │
    ├─────────► Datadog
    ├─────────► Grafana
    ├─────────► Elastic
    ├─────────► Azure Monitor
    ├─────────► AWS observability
    └─────────► custom platform

The architectural benefit is avoiding instrumentation that is tightly coupled to a single observability vendor.


14.13 What are the OpenTelemetry GenAI semantic conventions?

The GenAI semantic conventions are OpenTelemetry's standard names and meanings for model, token, conversation, agent, tool and retrieval telemetry.

They are worth naming explicitly in any enterprise architecture discussion.

Standard infrastructure telemetry understands:

HTTP request
database call
queue operation
RPC call

But GenAI introduces concepts such as:

model request
agent invocation
prompt
completion
token usage
tool call
retrieval
conversation
evaluation

OpenTelemetry has been developing GenAI-specific semantic conventions for these workloads. The current attributes cover concepts including GenAI provider/model information, conversations, workflows, retrieval documents, token usage, input/output messages and related AI telemetry. (OpenTelemetry)

For example, the OTel GenAI attribute registry includes concepts corresponding to:

gen_ai.provider.name
gen_ai.request.model
gen_ai.response.model
gen_ai.usage.input_tokens
gen_ai.usage.output_tokens
gen_ai.workflow.name

and retrieval/message telemetry. (OpenTelemetry)

Why semantic conventions matter

Without them:

Team A: prompt_tokens
Team B: inputTokenCount
Team C: llm.prompt.tokens

With a common convention:

gen_ai.usage.input_tokens

Now observability dashboards work across services.

Important security point

Do not indiscriminately record prompts and completions.

They can contain:

PII
PHI
credentials
contracts
source code
company secrets
customer data

The OTel specifications explicitly warn that GenAI input and output messages can contain sensitive/PII data. (OpenTelemetry)

So enterprise design should support:

capture disabled by default
redaction
masking
sampling
field-level filtering
encryption
RBAC
tenant isolation
retention policy

14.14 Trace hierarchy

A strong GenAI trace should expose logical AI operations, not merely HTTP calls.

For example:

TRACE: user request
│
└── Agent Run
    │
    ├── Policy Check
    │
    ├── Planning
    │
    ├── Retrieval
    │   ├── Query Rewrite
    │   ├── Vector Search
    │   └── Reranking
    │
    ├── Model Call
    │
    ├── Tool Call
    │   └── ERP / CRM / Database
    │
    ├── Model Call
    │
    ├── Verification
    │
    └── Human Approval

OpenTelemetry itself describes traces, metrics and events as important signals for GenAI telemetry, and recent OTel material illustrates model calls, tool invocations and token exchanges as parts of agent execution that observability must expose. (OpenTelemetry)


14.15 Agent run

The root AI span usually corresponds to:

one business-level agent invocation

Record attributes such as:

agent
agent_version
workflow
tenant
session
request type
start/end
status
termination reason
number of steps
number of model calls
number of tool calls
cost
quality result

Avoid adding unrestricted high-cardinality values such as raw user prompts into normal metric labels.


14.16 Model-call span

Track:

provider
requested model
returned model
temperature
max tokens
input tokens
output tokens
cache tokens if applicable
TTFT
total latency
retry
rate limit
finish reason
error
estimated cost

OpenTelemetry's current GenAI material specifically supports model identity and input/output token telemetry. (OpenTelemetry)


14.17 Retrieval span

You want to know:

query
query rewrite version
index
index version
filters
top_k
documents returned
similarity scores
retrieval latency
tenant/ACL filtering

Quality telemetry can add:

Recall@K
Precision@K
MRR
NDCG
context relevance
groundedness

The crucial point is:

Retrieval success cannot be inferred from HTTP 200.

The vector database may be perfectly healthy while retrieving useless documents.


14.18 Reranking span

Track:

reranker model
documents in
documents out
scores
latency
token consumption
cost

This allows questions such as:

Does reranking improve quality enough
to justify its extra latency and cost?

14.19 Tool-call span

For agents, tool telemetry is critical.

Record:

tool
tool version
operation
read/write
arguments metadata
authorization decision
latency
result status
retry
error
side effect
idempotency key

For sensitive arguments, capture:

hash / schema / metadata

instead of full values.

A tool invocation returning HTTP 200 could still be semantically wrong, so tool observability should include business outcome, not just transport status. A central tool gateway is the natural place to emit these spans consistently.


14.20 Policy-check span

Example:

policy = finance-agent-policy
version = 17
decision = denied
rule = purchase_limit
risk = high
latency = 11ms

This becomes invaluable for:

audit
AI governance
security investigations
regulatory review
RCA

14.21 Human-approval span

Record:

approval requested
approver role
requested timestamp
approved/rejected
decision timestamp
reason code

Usually avoid recording unnecessary personal information.

Human approval time should also normally be distinguished from machine-processing latency.

Otherwise:

agent latency = 3 hours

looks terrible although the agent itself took 2 seconds and waited 2h59m58s for a manager.


14.22 Verification span

Verification might include:

schema validation
grounding check
policy validation
citation validation
LLM verifier
deterministic rule
business-rule validation

Track:

verifier
version
pass/fail
score
reason

This is particularly useful for agent systems because successful generation is not equivalent to successful task completion.


14.23 The metrics hierarchy

I classify GenAI metrics into five layers.

LayerExamples
InfrastructureCPU, memory, GPU, queue depth
Applicationrequests, latency, errors
LLMTTFT, tokens, model errors, cost
AI qualitygroundedness, retrieval quality, safety
Businessresolution rate, conversion, cycle time
The higher you go, the closer the metric gets to business value.

An AI architect should resist having a dashboard that contains only:

CPU
requests
HTTP errors

while completely ignoring whether the AI is producing useful results.


14.24 Latency

Always measure percentile latency:

P50
P90
P95
P99

not just average.

And decompose:

Total latency
 =
gateway
+ orchestration
+ retrieval
+ reranking
+ model
+ tools
+ verification

Otherwise:

P95 = 14 seconds

doesn't tell you what to optimize.


14.25 What is TTFT (time to first token)?

TTFT is the time from sending a request to receiving the first streamed token or chunk of the response.

For streaming systems, TTFT can matter more to perceived latency than total generation duration.

Request ──────────────► first token ─────────────► complete
        <--- TTFT --->          generation

Users may tolerate:

8-second total generation

better when streaming starts in:

600 ms

than a five-second request that shows nothing until completion.

The OTel GenAI attributes include time-to-first-chunk semantics for streaming responses. (OpenTelemetry)


14.26 Tokens

Track at least:

input tokens
output tokens
tokens/request
tokens/session
tokens/tenant
tokens/user
tokens/application
tokens/model

Useful derived measures:

tokens / successful task
tokens / business transaction
tokens / resolved case

The last ones are more valuable than merely:

tokens/day

14.27 Token dashboards

A good token dashboard could expose:

TOTAL TOKENS
├── application
├── tenant
├── model
├── agent
├── workflow
├── prompt version
└── environment

INPUT / OUTPUT SPLIT

TOKENS PER REQUEST
P50 / P95 / P99

TOKEN TREND

TOP TOKEN-CONSUMING WORKFLOWS

FAILED REQUEST TOKEN WASTE

RETRY TOKEN WASTE

CACHE HIT RATE

TOKENS PER SUCCESSFUL TASK

Existing GenAI observability platforms commonly surface trace-level and aggregate token usage alongside model-call information, which reflects the operational importance of these measures. (Docs by LangChain)

Architect insight

High token consumption may indicate:

oversized RAG context
unnecessary conversation history
agent loops
poor prompt design
retrieval over-fetching
duplicated system prompts
failed retries
unnecessary verifier calls

So token telemetry is not only FinOps.

It can reveal architectural inefficiency.


14.28 Cost dashboards

Never stop at:

OpenAI bill = $82,000

You want chargeback/showback:

cost
├── tenant
├── business unit
├── application
├── agent
├── workflow
├── model
├── request
└── successful task

And separate:

LLM inference
embeddings
reranking
vector database
search
tool APIs
compute
storage
network
observability

Useful KPI:

Cost per successful task

instead of:

Cost per API call

For example:

Agent A
$0.04/request
60% success

effective cost/success ≈ $0.067


Agent B
$0.05/request
95% success

effective cost/success ≈ $0.053

The superficially more expensive agent is actually cheaper.

Commercial GenAI observability systems similarly aggregate token counts and derived costs at trace/model-call level. (Docs by LangChain) The budgeting and chargeback model these dashboards feed is laid out in AI FinOps: A Practical Framework to Control Enterprise AI Cost.


14.29 Errors

Differentiate:

Infrastructure error
Network error
Provider error
Rate limit
Timeout
Model refusal
Structured-output failure
Retrieval failure
Tool failure
Policy rejection
Verification failure
Agent-loop termination
Business failure

Because:

HTTP error rate = 0.01%

could coexist with:

business task failure = 18%

14.30 Retries

Retries are especially dangerous in GenAI systems because they amplify:

latency
cost
token consumption
provider throttling
side effects

Track:

retry rate
reason
retry success
tokens wasted
cost wasted
tool retries
model retries

A retry storm against a rate-limited model can become self-amplifying.


14.31 Tool failures

Track by:

tool
operation
error type
tenant
agent
version

Important derived metrics:

tool success rate
tool latency
tool retry rate
invalid argument rate
authorization rejection rate

A rising invalid-argument rate can indicate:

prompt drift
model migration
tool schema incompatibility
agent regression

rather than a tool problem.


14.32 Retrieval quality

Production retrieval monitoring should separate:

retrieval health

from:

retrieval quality

Health:

Vector DB available?
Latency good?
Queries successful?

Quality:

Did we retrieve useful evidence?

Measure where feasible:

context relevance
Recall@K
Precision@K
MRR
NDCG
citation correctness
groundedness
no-result rate
low-score rate

14.33 Task success

This is arguably the most important agent metric.

Don't define:

LLM returned answer = success

Define the business task.

Examples:

Customer query resolved
Purchase order successfully created
Invoice correctly processed
Incident resolved
Contract clause correctly identified
Supplier onboarding completed

Then monitor:

Task Success Rate =
successful business tasks / initiated tasks

This becomes the bridge between LLMOps and business operations.


14.34 What is drift in a GenAI system?

Drift is any change over time in the inputs, evidence, prompts, models or behaviour of a production AI system relative to its baseline.

Drift is broader in GenAI than traditional ML.

Five categories are worth separating, and I keep exactly this taxonomy: data, retrieval, prompt, model and behaviour drift.


Data drift

The input population changes.

Example:

Training / baseline:
80% English
20% Hindi

Production six months later:
40% English
60% Hindi

Or:

support queries become much longer
new product terminology appears
new customer segment arrives

Detection methods:

distribution monitoring
embedding-distribution comparison
topic clustering
length/language statistics
schema changes
feature distribution

14.35 What is retrieval drift?

Retrieval drift is when the retrieval system starts returning different or lower-quality evidence.

Possible causes:

corpus changed
new document types
embedding model changed
index changed
metadata degraded
ACL filtering changed
query distribution changed
stale documents accumulated

Monitor:

similarity distribution
no-result rate
top-K relevance
document age
source distribution
retrieval evaluation scores
citation rates

Each of those causes maps to a specific pipeline stage; RAG Architecture: The Full Pipeline and Where Each Stage Fails walks through them.

The one-line summary

A RAG application can drift even if the LLM never changes because the corpus and query distribution are dynamic parts of the model's effective environment.

14.36 Prompt drift

Prompt drift is a change in the prompt the model actually receives, whether someone edited the template or the dynamic context around it changed.

This can mean two things.

Explicit prompt drift

Someone changes:

system prompt
few-shot examples
prompt template
tool instructions

Solution:

prompt registry + version control

Effective prompt drift

The static prompt didn't change, but dynamic context did:

retrieved documents
conversation history
memory
tool descriptions
user context

So the effective prompt sent to the model changes.

This is one reason trace capture is so important.


14.37 Model drift

With external models, behaviour can change because of:

model upgrade
provider migration
endpoint configuration
quantization
routing
fine-tuning
regional deployment
provider-side changes

Mitigation:

pin versions where possible
maintain golden evaluations
shadow new models
canary
regression gates
monitor production quality

14.38 Behaviour drift

This is the highest-level form.

The system continues functioning technically but behaviour changes.

For example:

more refusals
more tool calls
longer reasoning
lower task completion
different tool selection
higher escalation
more verbose answers
less grounding

For agents this might be detected through:

tool-call distribution
step-count distribution
termination reasons
task success
verification failures
policy rejections
human escalation

14.39 Drift architecture

Conceptually:

Production telemetry
       │
       ▼
Feature / behaviour extraction
       │
       ├── input statistics
       ├── embedding distributions
       ├── retrieval scores
       ├── tool patterns
       ├── token patterns
       ├── evaluation results
       └── business outcomes
       │
       ▼
Baseline comparison
       │
       ▼
Threshold / statistical detector
       │
       ▼
Alert
       │
       ▼
Evaluation pipeline
       │
       ▼
Human investigation / remediation

Critical distinction:

Drift is a signal to investigate; it is not automatically proof of degradation.

A shift can be legitimate.


14.40 Production sampling

You generally cannot perform expensive evaluation on 100% of traffic.

So:

100% lightweight telemetry
        +
sampled rich traces
        +
sampled evaluation

Example strategy:

100%:
 latency
 tokens
 errors
 model
 tool status

10%:
 richer trace details

2%:
 automated LLM evaluation

100% of:
 errors
 policy violations
 low-confidence outputs
 high-value transactions

That last piece is risk-based sampling.

I prefer it over purely random sampling for enterprise AI.


14.41 Feedback loops

Production feedback can come from:

Explicit
    thumbs up/down
    rating
    user comment

Implicit
    user retries
    reformulation
    abandonment
    escalation
    answer copied
    task completion

Operational
    verifier failure
    policy rejection
    tool failure

Business
    conversion
    resolution
    revenue
    cycle time

Architecture:

Production interaction
        ↓
Telemetry + feedback
        ↓
Feedback store
        ↓
Curated difficult cases
        ↓
Evaluation dataset
        ↓
Prompt/model/agent improvements
        ↓
Regression evaluation
        ↓
Deployment

This creates the continuous AI improvement loop.

Modern LLM evaluation systems increasingly connect evaluations directly to traces so production feedback and offline evaluation can share the same observability lineage. (MLflow AI Platform)


14.42 Incident management

AI incidents differ from traditional outages.

Traditional:

service unavailable
database down
latency high

AI incidents may be:

hallucination spike
unsafe output
PII disclosure
wrong tool execution
retrieval corruption
prompt injection
agent loop
cost explosion
model degradation
wrong business decision

So classify incidents across:

Availability
Performance
Quality
Security
Safety
Cost
Compliance
Business

14.43 AI incident response

A mature response flow is:

DETECT
  ↓
CONTAIN
  ↓
IDENTIFY BLAST RADIUS
  ↓
PRESERVE EVIDENCE
  ↓
MITIGATE
  ↓
ROLL BACK / DISABLE
  ↓
ROOT CAUSE
  ↓
RE-EVALUATE
  ↓
RELEASE
  ↓
POST-INCIDENT REVIEW

Potential containment actions:

disable tool
disable memory
switch model
rollback prompt
rollback retriever
disable agent
force human approval
reduce autonomy
route to safe workflow
block affected tenant/use case

This is why feature flags and kill switches belong in AI architecture.


14.44 Rollback

The hardest thing about LLM rollback is that there may not be one thing to roll back.

The failure may have come from:

prompt
model
retriever
embedding
index
tool
agent graph
policy
memory
application code

Therefore each should be independently versioned.

You want:

Current:
App 14
Prompt 27
Model B
Index 42
Policy 9

Incident caused by Prompt 27

Rollback:
Prompt 27 → Prompt 26

rather than redeploying the entire platform.


14.45 Root-cause analysis

Imagine:

“Customer-service task success dropped from 91% to 73%.”

A weak investigation checks application logs.

A strong AI RCA follows the trace hierarchy:

1. Traffic changed?
        ↓
2. Application errors?
        ↓
3. Model latency/errors?
        ↓
4. Model/version changed?
        ↓
5. Prompt changed?
        ↓
6. Retrieval changed?
        ↓
7. Corpus/index changed?
        ↓
8. Tool behaviour changed?
        ↓
9. Policy changed?
        ↓
10. Evaluation/judge changed?
        ↓
11. Business population changed?

Then correlate the timeline:

09:00 index deployment v42
10:00 retrieval relevance ↓
10:15 groundedness ↓
10:30 task success ↓

Now RCA becomes evidence-driven.


14.46 AIOps in this architecture

Don't confuse LLMOps with AIOps.

LLMOps

Operating the AI application itself:

models
prompts
RAG
agents
evaluation
deployment
observability

AIOps

Using analytics/AI to operate the technology environment:

anomaly detection
event correlation
incident prediction
root-cause assistance
capacity forecasting
automated remediation

Example:

Telemetry:
LLM latency ↑
rate limit ↑
retry rate ↑
token spend ↑

AIOps correlation:
"Provider rate limiting appears to be triggering retry
amplification causing both latency and cost escalation."

Recommended action:
route 30% traffic to fallback model.

At greater maturity:

Detect
  ↓
Diagnose
  ↓
Recommend
  ↓
Human approve
  ↓
Remediate

For lower-risk situations:

Detect → Diagnose → Automatically remediate

Always constrain autonomous remediation with policy and blast-radius limits.


14.47 How is LLM observability different from traditional observability?

This distinction is worth memorising.

TraditionalGenAI
HTTP successTask success
CPUTokens
Response latencyTTFT + generation latency
ExceptionHallucination
DB queryRetrieval
RPCTool call
Build versionPrompt + model + dataset versions
Error rateGroundedness / safety
Request traceAgent trajectory
Infrastructure driftBehaviour drift
A system can therefore be 100% technically healthy while being AI-functionally broken.

That is the single most important point in this domain.


14.48 Enterprise multi-tenant observability

At enterprise-platform level, every trace should carry context such as:

tenant_id
application_id
agent_id
environment
region
deployment_version

Potentially:

business_unit
use_case
risk_class

But be careful about:

user IDs
document IDs
raw prompts
customer names

because of privacy and metrics-cardinality problems.

Access should enforce:

tenant A cannot query tenant B traces

application owners → own applications
platform team      → platform telemetry
security           → security traces
audit              → immutable audit records

Observability itself becomes a multi-tenant platform capability. The same tenant-context propagation that protects data paths, described in Tenant Context in Multi-Tenant Microservices, has to reach the telemetry path too.


14.49 Do not put everything into telemetry

A common architecture error is:

“We'll log every prompt, completion, document and tool response.”

That creates enormous:

privacy risk
security risk
storage cost
compliance risk
observability cost

Instead classify telemetry:

METADATA
Safe for widespread capture

CONTENT
Restricted / sampled

SENSITIVE CONTENT
Redacted or prohibited

AUDIT EVENTS
Separate immutable retention

OpenTelemetry's GenAI guidance explicitly treats prompt/output content as potentially sensitive, reinforcing this separation. (OpenTelemetry)


14.50 The dashboard I would build

An enterprise AI operations dashboard might look conceptually like:

──────────────── AI PLATFORM HEALTH ────────────────

REQUESTS          TASK SUCCESS      P95 LATENCY
2.4M/day              94.1%             3.8s

LLM ERRORS        TOOL ERRORS       RETRIEVAL QUALITY
0.12%                0.8%               91%

TOKENS/DAY        COST/DAY          COST/SUCCESS
780M                $18,200              $0.044


MODEL
Model A    72%
Model B    24%
Fallback    4%

QUALITY
Groundedness        94%
Correctness         92%
Safety              99.8%
Retrieval relevance 91%

AGENTS
Support agent        96% success
Procurement agent    92%
Finance agent        89%   ⚠

DRIFT
Input drift       normal
Retrieval drift   warning
Behaviour drift   warning

OPERATIONS
Model retries     ↑ 18%
Tool failures     normal
Policy rejects    ↑ 7%

Tools such as LangSmith currently expose similar categories including trace volume, latency/error rates, LLM calls, token/cost usage, tool errors/latency and feedback scores, although an enterprise architecture should retain vendor-neutral telemetry underneath where practical. (Docs by LangChain)


14.51 Observability pipeline design

For a large enterprise-scale architecture, I draw this:

                  APPLICATIONS
                       │
        ┌──────────────┼──────────────┐
        │              │              │
       RAG           Agents        LLM APIs
        │              │              │
        └──────── OTel SDK ───────────┘
                       │
                       ▼
                OTel Collectors
                       │
             sampling / filtering
             PII redaction
             enrichment
             batching
                       │
      ┌────────────────┼───────────────────┐
      │                │                   │
      ▼                ▼                   ▼
 Trace Backend     Metrics Backend     Log Backend
      │                │                   │
      └────────────────┼───────────────────┘
                       ▼
                AI Observability
                       │
       ┌───────────────┼────────────────┐
       │               │                │
   Dashboards        Alerts          Evaluation
       │               │                │
       └───────────────┼────────────────┘
                       ▼
                 Incident / Ops
                       │
                       ▼
             Continuous Improvement

This keeps application instrumentation independent of the backend provider.


14.52 What SLOs does a GenAI system need?

A GenAI SLO set covers reliability, quality, safety and economics together, not availability and latency alone.

Traditional SLO:

99.9% availability
P95 < 2s

GenAI needs multidimensional SLOs.

For example:

Availability           ≥ 99.9%
P95 TTFT                < 1.5s
P95 total latency       < 8s
Task success            ≥ 93%
Groundedness            ≥ 95%
Tool success            ≥ 99%
Critical safety failure < 0.01%
Cost/task               < $0.06

Notice:

Reliability + quality + safety + economics.

That is the correct mental model.


14.53 The architect-level trade-off

You will often face:

QUALITY
   ▲
   │
   │
COST ───────── LATENCY

Adding:

reranker
verifier
larger model
more retrieved context
multiple agent workers

may increase quality but also:

latency ↑
tokens ↑
cost ↑
failure surface ↑

So LLMOps gives you the data to make that trade-off quantitatively.


14.54 Failure modes you should know

FailureDetectionMitigation
Prompt regressionevaluation ↓rollback prompt
Provider model regressionquality ↓model rollback/routing
Token explosiontokens/request ↑context limits
Agent loopsteps/request ↑max iterations
Retrieval driftrelevance ↓reindex/retriever fix
Tool degradationtool errors ↑circuit breaker
Provider throttling429/retry ↑backoff/routing
Cost spike$/task ↑budget controls
PII in telemetryaudit/DLPredact/filter
Evaluation driftevaluator distribution changesversion evaluators
Behaviour drifttool/trajectory distributions changeinvestigate/evaluate
RAG stalenesssource age ↑freshness pipeline
---

14.55 FAQ: How would you design observability for an enterprise GenAI platform?

I design GenAI observability at three levels: platform health, AI behaviour and business outcome. At the AI level I instrument the entire execution graph (agent runs, model calls, retrieval, reranking, tool calls, policy checks and verification) with OpenTelemetry so the instrumentation stays vendor-neutral and follows the GenAI semantic conventions, and every trace carries lineage for the application, model, prompt, retriever, policy and deployment versions so any production result can be reproduced and root-caused. I monitor latency including TTFT, token consumption, cost, model and tool errors, retries and retrieval quality, but I also measure task success, groundedness, safety and business outcomes, because an AI application can return HTTP 200 and still fail functionally. In production I use risk-based sampling for expensive evaluations, monitor data, retrieval, model and behaviour drift, and feed difficult production cases back into the evaluation dataset. Finally I make prompts, models, retrieval configuration and agent versions independently deployable and rollbackable, with canaries, feature flags and kill switches, which gives both observability and operational control.

What separates this from a pure MLOps answer is the business-outcome layer and independent rollback of every component.


14.56 FAQ: Why use OpenTelemetry for GenAI?

GenAI systems are distributed systems: a single request can involve the application, agent orchestrator, retriever, vector database, model gateway, external LLMs and enterprise tools. OpenTelemetry gives me a common vendor-neutral trace context across those components. The GenAI semantic conventions extend that model with concepts such as model information, token usage, conversations and AI-specific operations. The result is that I can correlate traditional infrastructure telemetry with the actual AI execution path instead of maintaining a separate black-box AI monitoring stack.

OpenTelemetry describes this common semantic scheme as enabling easier correlation and consumption of telemetry across codebases, libraries and platforms. (OpenTelemetry)


14.57 FAQ: How do you detect an LLM production issue?

I work from the business outcome down: task success first, then quality, behaviour, the LLM layer, RAG, tools, configuration, traffic and infrastructure. The final check is the measurement system itself, because a changed judge or evaluator can look exactly like a regression. The question that is easiest to forget is: did the system regress, or did the evaluator change?

The full sequence:

1. Business outcome
   Did task success change?

2. Quality
   Did correctness/groundedness/safety change?

3. Behaviour
   Did trajectories/tool usage/step count change?

4. LLM
   Did model/version/token/latency/error patterns change?

5. RAG
   Did retrieval relevance or corpus change?

6. Tools
   Did downstream APIs change?

7. Configuration
   Prompt / agent / policy / feature flag change?

8. Traffic
   Did the user/query population change?

9. Infrastructure
   CPU/network/database/provider problems?

10. Evaluation
   Did the measurement system itself change?

14.58 FAQ: What is the most important LLMOps metric?

Not tokens, and not latency. The most important metric is task success or business outcome, supported by quality, reliability, latency and cost metrics as guardrails. That keeps LLMOps tied to business value rather than model activity.

For example:

North Star:
Successful procurement requests

Guardrails:
≥95% groundedness
≥99.9% availability
P95 < 8 seconds
<$0.08/task
zero critical-policy violations

14.59 FAQ: Should you log every prompt and completion?

No. Prompts and completions can carry PII, PHI, credentials, contracts, source code and customer data, and the OpenTelemetry GenAI specification itself flags them as potentially sensitive. I capture metadata broadly, keep content capture off by default and sampled where it is needed, redact or prohibit sensitive content, and route audit events to separate immutable retention. That keeps traces useful for RCA without turning the observability platform into the largest unprotected copy of customer data.


14.60 FAQ: How do you roll back an LLM application safely?

I version every component that can change behaviour independently: prompt, model, retriever, embedding, index, tool, agent graph, policy and memory. When an incident traces to one component, I roll back only that component, for example moving the production prompt alias from v27 to v26, rather than redeploying the whole platform. Feature flags and kill switches give me containment options such as disabling a tool, forcing human approval or switching models while the root cause is confirmed.


14.61 FAQ: What is the difference between LLMOps and AIOps?

LLMOps is operating the AI application itself: models, prompts, RAG, agents, evaluation, deployment and observability. AIOps is using analytics and AI to operate the technology environment, through anomaly detection, event correlation, incident prediction, root-cause assistance, capacity forecasting and automated remediation. The two meet when AIOps correlates LLM telemetry, for example spotting that provider rate limiting is driving retry amplification, and I always constrain any autonomous remediation with policy and blast-radius limits.


14.62 One architecture principle worth memorising

The entire topic collapses into:

OBSERVE
      ↓
UNDERSTAND
      ↓
EVALUATE
      ↓
CONTROL
      ↓
IMPROVE

Or, more concretely:

Version everything.
Trace everything important.
Measure technical + AI + business outcomes.
Evaluate continuously.
Detect drift.
Rollback independently.
Feed production learning back into evaluation.

14.63 What you should know cold

In short:

  • Traditional observability shows whether the system runs; GenAI observability must show whether it behaves correctly.
  • Version every behaviour-changing artifact (prompt, model, corpus, retriever, policy, evaluator) and carry that lineage on every trace.
  • Instrument the logical execution graph with OpenTelemetry and the GenAI semantic conventions, without capturing raw content by default.
  • Measure task success and cost per successful task, with quality, latency, safety and cost as guardrails.
  • Treat drift as a signal to investigate, sample by risk, and roll back components independently.

The full list:

  1. LLMOps: lifecycle management of LLM applications.
  2. Prompt versioning: prompts are production artifacts.
  3. Model versioning: track the exact model and configuration.
  4. Dataset versioning: reproducible training, evaluation and RAG data.
  5. Evaluation versioning: version judges, judge prompts and thresholds.
  6. Deployment versioning: an immutable, complete AI configuration.
  7. Feature flags: canary, A/B and kill switch.
  8. Experiment tracking: configuration → result lineage.
  9. Traceability: reconstruct one complete request.
  10. Distributed tracing: trace across all services.
  11. OpenTelemetry: the vendor-neutral telemetry standard and ecosystem.
  12. GenAI semantic conventions: a common AI telemetry vocabulary.
  13. Agent span: the entire agent execution.
  14. Model span: one LLM invocation.
  15. Retrieval span: one RAG retrieval.
  16. Tool span: one external action or API call.
  17. TTFT: request → first streamed token or chunk.
  18. Tokens: a usage, cost and efficiency indicator.
  19. Cost per task is a better measure than cost per request.
  20. Task success is the key agent KPI.
  21. Data drift: the input distribution changes.
  22. Retrieval drift: evidence retrieval changes.
  23. Prompt drift: the prompt or effective context changes.
  24. Model drift: model behaviour or version changes.
  25. Behaviour drift: agent behaviour changes.
  26. Production sampling: evaluate representative and risky traces.
  27. Feedback loop: production → evaluation dataset → improvement.
  28. Incident management: detect, contain, RCA, recover.
  29. Rollback: independent component rollback.
  30. AIOps: AI-assisted IT operations.

The one sentence to remember

“Traditional observability tells me whether the application is running; GenAI observability must tell me whether it is behaving correctly, and every production response should be traceable to its model, prompt, retrieval configuration, data/index, agent, policies and deployment version.”

And the optimisation rule that follows from it: “I don't optimise an LLM platform for tokens or latency independently; I optimise cost, latency and reliability subject to the required task-success, quality and safety thresholds.”

That is the architect framing for LLMOps, AIOps and Observability.


Part of the series

The Enterprise AI Architect's Handbook
  1. 1.The Enterprise AI Architect Roadmap: The 29 Domains the Role Actually Owns
  2. 2.The AI Architect Operating Model: Turning a Business Objective into an Architecture
  3. 3.LLM Fundamentals for Architects: Tokens, Context, Latency, Throughput and Cost
  4. 4.Prompt and Context Engineering as an Architectural Concern
  5. 5.RAG Architecture: The Full Pipeline and Where Each Stage Fails
  6. 6.Knowledge Architecture: Ontologies, Entity Resolution and Graph Retrieval
  7. 7.Agent Architecture: Loops, Planning, Verification and Termination
  8. 8.Agent State and Memory Architecture: Scoping, Retention and Provenance
  9. 9.Multi-Agent Systems: When They Help, and How They Fail
  10. 10.Agent Orchestration: Frameworks, Durable Execution and Framework-Independent Design
  11. 11.MCP Architecture and the Enterprise Tool Gateway
  12. 12.Model Strategy: Selection, Gateways, Routing and Fallbacks
  13. 13.Fine-Tuning, RAG or Prompting: How an Architect Decides
  14. 14.Evaluating LLM, RAG and Agent Systems: Metrics, Judges and Quality Gates
  15. 15.LLMOps and Observability: Tracing, Metrics, Drift and Feedback Loops← you are here
  16. 16.AI Security: The Full Threat and Control Map for Architects
  17. 17.Responsible AI, Privacy and Governance as Architecture, Not Paperwork
  18. 18.Software Engineering for AI Platforms: The Non-Negotiable Baselinecoming soon
  19. 19.Cloud Architecture for AI Workloads: Isolation, Identity, Networking and Servingcoming soon
  20. 20.Containers, Infrastructure as Code and Delivery for AI Systemscoming soon
  21. 21.Cost and Performance Architecture: Designing for Cost per Successful Taskcoming soon
  22. 22.Reliability and Resilience: The Twenty Failure Modes of AI Systemscoming soon
  23. 23.Enterprise AI Platform Architecture: Control Plane and Runtime Planecoming soon
  24. 24.Production and Launch Readiness for AI Systemscoming soon
  25. 25.Domain Architecture: Applying the Model to a Real Business Functioncoming soon
  26. 26.AI System Design Practice: Fifteen Problems and How to Approach Themcoming soon
  27. 27.Architecture Artefacts: The Diagrams an AI Architect Must Be Able to Drawcoming soon
  28. 28.Structured Answers: System Design, Trade-offs, Incidents and Reviewscoming soon
  29. 29.Experience Narratives: The Stories an Architect Must Be Able to Tellcoming soon
  30. 30.Architecture Leadership and Technical Strategycoming soon
View full series →
AISeriesOctober 3, 2026
Share
Aakash Ahuja

Aakash Ahuja

Enterprise AI, Cybersecurity & Platform Engineering

Aakash writes about secure AI agents, microservices architecture, enterprise platforms, and production engineering. He has 20+ years of experience building and operating software systems across banking, cloud, cybersecurity, AI, and enterprise workflow automation. He is Director of Technology at itmtb Technologies and teaches AI, Big Data, and Reinforcement Learning at top institutes in India.