A diagnostic evaluation framework for industrial tool-use agents

ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents

Wei Chen1,*Peilun Zhou2,*,‡Zhaoyu Hu2Jiajun Chai2Zhongni Hou2Yufei Zhang2Derong Xu1Guojun Yin2,†Wei Lin2,†Zhi Zheng1,†Tong Xu1,†

1 State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China 2 Meituan

* Equal Contribution‡ Project Leader† Corresponding Author

Abstract

Large language model agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation must expose capability deficiencies, inform iteration priorities, and assess the effects of interventions.

Industrial agent service unfolds through the iterative trajectory of a current request and through continued interaction with a user. Reducing both to a final outcome can obscure where a deficiency emerges during execution and whether later service remains aligned with earlier context.

ATLAS is a dual-horizon diagnostic evaluation framework: trajectory-wise diagnostic signals relate observed deficiencies to execution locations and capability concerns, while user-wise signals assess whether service remains responsive across continued user interaction. Executable signals support calibrated evaluation, policy feedback, and reassessment.

On Meituan Xiaotuan, offline and online studies test this chain from diagnostic-signal fidelity to live-service outcomes.

Why evaluation must diagnose

Outcomes alone cannot guide agent iteration.

Industrial tool-use agents interpret user goals, invoke external tools, and complete real-world tasks in user-facing products and business workflows. Their behavior depends on dynamic business evidence and, across continued use, on a user’s evolving context.

Development teams have finite capacity to improve and deploy capabilities. They need evidence that exposes where a service falls short and makes scenario-specific priorities explicit—not only an aggregate score for one system version.

What outcome reporting hidesIt cannot show whether a process failed during execution, whether a decision used current business evidence, or whether later service still responds to the evolving interaction.

The iteration mandate

Diagnostic evidence must support three linked production decisions.

  1. 01What should improve?
  2. 02What should be prioritized?
  3. 03Did the intervention work?

Xiaotuan, Meituan’s AI assistant for local services, makes the requirement concrete. To turn an open-ended request into a decision-supporting recommendation, it iteratively retrieves and compares business-grounded information, reconciles relevant constraints, and updates later actions from intermediate observations and tool feedback. This workflow creates three evaluation challenges that outcome reporting alone cannot expose:

  • Challenge 01 — Errors can accumulate through a trajectory. A deviation in interpreting a request or selecting an action can compound through later tool use, even if the final response appears acceptable.
  • Challenge 02 — A recommendation must reflect live business conditions. Search results, merchant supply, and availability change with the current business state. Because open-ended user needs can admit several reasonable paths, an answer that sounds acceptable may still fail to provide an accurate, timely, and actionable recommendation.
  • Challenge 03 — Request-level success cannot establish sustained service quality. An ongoing interaction is more than a sequence of isolated requests. Each later turn may appear reasonable on its own while overlooking an earlier reference, constraint, or correction; over time, these locally plausible decisions can drift away from the user’s evolving goal.
A real Xiaotuan local-services interaction. The assistant iteratively retrieves and compares business-grounded information before synthesizing a recommendation.

Why dual-horizonThe first two challenges concern the current request trajectory: ATLAS’s within-turn horizon inspects the execution and its use of current business evidence. The third concerns continued service: its across-turn horizon evaluates user-relevant state across complete lifecycles.

ATLAS’s central design

Two horizons for diagnosing agent service.

Industrial agent service unfolds across two connected temporal scopes. ATLAS does not collapse them into one score: it retains distinct evidence for the complete trajectory triggered by a current request and for the continued interaction with one user.

Why the separation mattersWithin-turn evaluation localizes a capability deficiency in an execution. Across-turn evaluation asks whether subsequent service remains responsive as user-relevant context evolves.

Horizon 01 · Within-turn / trajectory-wise

Localize a gap in one request.

Examine the complete trajectory initiated by a request: thinking and reflection, tool or skill execution, updates after tool feedback, and response generation.

Diagnostic questionWhere does a capability deficiency arise in this execution?

Horizon 02 · Across-turn / user-wise

Assess responsiveness over time.

Examine whether referents, intent, constraints, and explicit corrections are carried forward, updated, and used across complete lifecycles.

Diagnostic questionDoes later service still respond to the user’s evolving context?

Dual-horizon evaluation frameworkTwo connected temporal scopes
The framework applies the two diagnostic horizons to agent service: trajectory-wise evidence for a current request and user-wise evidence across successive complete lifecycles.

Within-turn diagnostic structure

How the within-turn horizon localizes a gap.

Once a current-request trajectory is in scope, evaluation must retain both where behavior was observed and how it fell short of the task requirement. This structure expands only the within-turn horizon along those two complementary axes.

Across-turn stays outside this matrix by design: it evaluates relations among complete lifecycles and the evolving user context.

Execution location
Thinking & Reflection · Tool & Skill Execution · Response Generation
Capability concern
Relevance · Factuality · Timeliness · Reliability · Intent & Planning
Guardrail
Norms & Compliance covers safety, privacy, compliance, and structural validity.
Within-turn diagnostic structureExecution location × capability concern

From diagnosis to iteration

From diagnostic signals to feedback.

The structure becomes operational only when each signal specifies the behavior to assess, the evidence it may use, and the boundary for a deficient outcome. ATLAS then calibrates those decisions, scales selected signals where needed, and reuses them as policy feedback before reassessing the updated policy in the same framework.

The semantic threadA signal’s target, evidence scope, and decision boundary are established before calibration and retained when the resulting policy is evaluated again.

  1. 01
    Specify the signal

    For each behavior, define the target, admissible evidence, decision boundary, and reporting semantics. Use rules when the boundary is programmatic and an LLM judge when it requires semantic assessment.

  2. 02
    Calibrate LLM-based signals

    Evaluate each signal-specific judge against high-confidence references from real business logs. Refine evidence binding, exemptions, and decision criteria until it reproduces the intended boundary reliably.

  3. 03
    Make selected signals efficient

    Selected compatible LLM-based signals can be distilled into lower-latency, lower-cost diagnostic models that retain signal-specific decision behavior.

  4. 04
    Optimize, then reevaluate

    Selected calibrated signals provide multidimensional feedback for policy optimization. Assess the updated policy again with trajectory-wise and user-wise evidence.

A diagnostic result is not an automatic product decision. It identifies a candidate capability deficiency; product context and deployment constraints still determine priority and intervention.

The diagnostic structure guides signal construction and calibration; selected calibrated signals support policy feedback and subsequent reassessment.

Companion resource · Public signal inventory

Inspect individual signals.

The method above defines how signals are specified, calibrated, scaled, and reused for iteration. The public Signal Explorer exposes the trajectory-wise and user-wise signals that instantiate that method, each with a target, evidence scope, decision boundary, and reporting semantics.

Use it as a referenceMove from the dual-horizon structure on this page to individual signal definitions and the evidence each decision may use.

Open the Signal Explorer

Browse the inventory for signal-level detail, then continue to the validation evidence below.

Preview of the ATLAS Signal Explorer, which maps its trajectory-wise and user-wise diagnostic signals.
Open the public ATLAS Signal Explorer ↗

Validation on Meituan Xiaotuan

A validation chain from signal decisions to live service.

The experiments test a connected chain: whether diagnostic signals implement their intended decisions, whether selected signals can run efficiently at scale, whether their feedback improves service under deployment-aligned replay, and whether those gains transfer to live traffic.

  1. 01Judge interfacesCan signal boundaries be reproduced?
  2. 02Efficient modelsCan selected signals run at scale?
  3. 03Offline replayDoes feedback improve replayed service?
  4. 04Online A/B & auditDo gains reach live traffic?

Experimental setting

All studies use mutually query-disjoint data from real Xiaotuan production traffic. High-confidence real-log references calibrate and evaluate 41 LLM-based signals; separate data train 17 efficient diagnostic models. Policy evaluation replays about 2,000 queries through production-facing tools with live business data, while optimization uses a separate 10,000-query set.

Study 01 · Judge interfaces

Can an LLM judge reproduce a signal’s intended decision boundary?

Purpose
A diagnostic signal is useful only if its semantic judgment implements the intended boundary.
How
Four judge interfaces share the same DeepSeek-V4-Pro backbone and are evaluated with F1 on high-confidence references for 41 LLM-based signals from real Xiaotuan logs.
Average F1 across comparable LLM-based diagnostic signals.
MethodRel.Fact.Time.Reliab.IntentN&CUser-wiseOverall
Direct Judge0.8310.8420.6910.8110.7680.8130.7840.800
Static Rubric0.8430.8530.8490.7750.8110.8760.7610.826
Curated Rubric0.8850.9050.8380.8080.8430.9320.8800.877
ATLAS Judge0.9390.9560.9610.9720.9370.9840.9530.952

FindingATLAS Judge is strongest in every reported group, reaching 0.952 overall F1; calibrated evidence binding and boundary control faithfully implement the diagnostic decisions.

ColumnsRel. — Relevance · Fact. — Factuality · Time. — Timeliness · Reliab. — Reliability · N&C — Norms & Compliance

The comparison covers LLM-based signals; rule-based signals follow their programmatic specifications.

Study 02 · Efficient models

Can selected signals retain their decision behavior in a compact model?

Purpose
High-capacity LLM judges are costly to invoke repeatedly at scale.
How
We distill 17 selected LLM-based signals onto Qwen3.5-9B and compare their F1 against the same high-confidence references, the untuned 9B backbone, and larger open baselines.
Group-average F1 (%) for 17 selected diagnostic signals.
Signal group DeepSeek-V4-Pro Qwen3.5-9B Qwen3.6-27B Qwen3.6-35B-A3B Ours Δ vs 9B
Relevance95.276.990.888.292.7+15.8
Factuality95.981.484.182.394.6+13.2
Reliability97.278.095.585.695.5+17.5
Norms & Compliance98.487.396.284.496.7+9.4
Overall96.481.090.384.994.6+13.6

FindingThe distilled 9B model reaches 94.6% overall F1, a +13.6-point gain over its untuned 9B backbone, showing that selected calibrated decisions can be deployed more efficiently.

DeepSeek-V4-Pro is the high-capacity reference interface. Δ is the absolute gain over Qwen3.5-9B.

Studies 03–04 · Policy outcomes

Does diagnostic feedback improve service?

Offline replay covers three within-turn execution dimensions, the Norms & Compliance guardrail, and the user-wise layer.

Study 03 · Offline replay

Does feedback improve replayed service?

Purpose
Test whether selected calibrated signals provide useful policy feedback beyond improving the final response alone.
How
Real Xiaotuan queries are replayed to each policy; production-facing tools resolve merchants, POIs, and supply from live business data at execution time.

Finding. The ATLAS-optimized policy has the highest reported mean score in every shown group, including Tool & Skill Execution and the user-wise layer, while Norms & Compliance remains high.

Study 04 · Online A/B & audit · 10% treatment

Do replay gains transfer to production?

Purpose
Test whether gains observed under deployment-aligned replay transfer to production service.
How
The optimized policy is compared concurrently with the deployed version from June 27 to June 29 on a fixed 10% slice of live traffic. Every reported product metric moves in the favorable direction.

Δ values are relative changes versus the deployed version (%)

AI message read-through rate
+0.92
Per-user dwell time
+1.26
AI search query volume
+0.76
Effective QV
+0.50
Session follow-up rate
+6.9
Paid GTV
+7.32
5-second session exit rate
−0.31

Companion audit · 1,000 responses

Do audits confirm outcomes?

Purpose
Check whether sampled response quality shifts in the same direction as product outcomes.
How
A 1,000-response audit from the same experiment window measures response–supply relevance and hallucination rates.

Δ values are absolute rate changes (pp)

Response–supply relevance
+6.2
P0 hallucination rate
−6.9
ID hallucination rate
−2.0

Interpretation note: Session follow-up is a session-level indicator of continued interaction, not a single user-wise signal.

Conclusion

Evaluation as evidence for targeted iteration.

ATLAS is a dual-horizon diagnostic framework for industrial tool-use agents. Trajectory-wise signals localize capability gaps in a current execution; user-wise signals assess whether subsequent service remains responsive as interaction context evolves.

Fine-grained, evidence-bound signals make both views executable. Selected calibrated signals support scalable evaluation and multidimensional policy feedback, while the same framework reassesses the resulting policy.

On Xiaotuan, this chain is validated from judge fidelity and efficient diagnostic models to offline replay, online product outcomes, and sampled human-audit quality.