Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 14 additions & 1 deletion docs/metrics/turn_taking.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ Code-based metric (no LLM) that scores each user→assistant transition on a con
- `latency_assistant_turns` — per-turn latency (`first_asst_start - last_user_end`) in seconds. Drives the latency curve.
- `assistant_interrupted_turns` / `user_interrupted_turns` — turn-level interruption flags set by the processor.
- `conversation_trace` — used to detect which turns contain a tool call (`type == "tool_call"`). Tool-call turns get a more lenient upper end of the latency curve and a higher `late` threshold, since tool execution adds inherent latency.
- `output_dir` — read directly (not via `conversation_trace`) to load `audit_log.json` for the `pretoolspeech_rate` sub-metric (see below).

## Per-turn Score

Expand Down Expand Up @@ -189,7 +190,19 @@ Sub-metric aggregation reads the raw per-turn values from `per_turn_evidence` ra
| `user_interruption.mean_yield_ms` | no | rate > 0 | Arithmetic mean of `yield_ms` across user-interrupt turns. |
| `user_interruption.mean_yield_score` | yes | rate > 0 | Mean of the per-turn yield scores that feed the main score. |

Rate sub-metrics are emitted as `normalized_score` (they already live on `[0, 1]`). Raw-ms sub-metrics have `normalized_score = None` so they don't corrupt cross-metric averages.
**Pre-tool-speech lead-in rate**

| Key | Normalized? | When present | Meaning |
| --- | --- | --- | --- |
| `pretoolspeech_rate` | yes | at least one tool-call group and `audit_log.json` is readable | Fraction of tool-call groups preceded by a non-empty assistant utterance since the previous group ended (or since the conversation started). Measures how often the agent gives a spoken lead-in (per `agent.pre_tool_speech` in `configs/prompts/simulation.yaml`) before invoking a tool. |

Computed by `_compute_pre_tool_speech_groups`, which reads `audit_log.json` directly from `context.output_dir` **instead of** `conversation_trace`. A "tool-call group" is a maximal contiguous run of `tool_call`/`tool_response` entries in the raw audit-log transcript (sorted by `timestamp`) — consecutive tool calls fired off one lead-in (e.g. two tools called back-to-back with no intervening speech) count as a single group, matching the prompt's "one lead-in per batch" instruction. Assistant speech only counts if it has non-empty, non-whitespace content; a `user` event resets the "spoke since last group" flag so a lead-in from a previous turn can't be credited to a later turn's tool calls. Event types other than `assistant`/`user`/`tool_call`/`tool_response` (e.g. `llm_call`) are ignored rather than treated as group-breaking — see the regression test guarding this.

**Why not `conversation_trace`?** For S2S, `conversation_trace` assistant entries come from the user simulator's STT transcription of the spoken audio, timestamped only after the audio finishes playing — not the audit log's true (early) LLM-generation timestamp used to build the trace for cascade/audio-LLM. That transcription can sort *after* a subsequent tool call in the timestamp-ordered trace even when the agent actually spoke first (confirmed against a real S2S transcript), so trace order can't be trusted for this signal on S2S. Reading `audit_log.json` directly sidesteps the issue and gives one code path across cascade, S2S, and audio-LLM.

`pretoolspeech_rate` is omitted (not emitted) when there are no tool-call groups, or when `audit_log.json` is missing/unreadable/empty at `context.output_dir` — treated as "unknown", not "no lead-ins".

Rate sub-metrics are emitted as `normalized_score` (they already live on `[0, 1]`). Raw-ms/count sub-metrics have `normalized_score = None` so they don't corrupt cross-metric averages.

## Tunable Constants

Expand Down
5 changes: 5 additions & 0 deletions src/eva/assistant/agentic/audio_llm_system.py
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@ def __init__(
audit_log: AuditLog,
alm_client: BaseALMClient,
output_dir: Path | None = None,
pre_tool_speech: str = "off",
llm_streaming: bool = False,
full_audio_context: bool = False,
):
Expand All @@ -49,6 +50,7 @@ def __init__(
audit_log=audit_log,
llm_client=alm_client,
output_dir=output_dir,
pre_tool_speech=pre_tool_speech,
llm_streaming=llm_streaming,
)
self.alm_client: BaseALMClient = alm_client
Expand All @@ -64,6 +66,9 @@ def __init__(
agent_instructions=agent.instructions,
datetime=current_date_time,
)
# Reuse the shared pre-tool lead-in prompt appended to the system prompt.
if self.pre_tool_speech == "auto":

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you make this case insensitive please? Auto, AUTO, auto

self.system_prompt += "\n\n" + self.prompt_manager.get_prompt("agent.pre_tool_speech")

# Per-turn audio history: list of (audio_bytes, sample_rate)
self._turn_audio_history: list[tuple[bytes, int]] = []
Expand Down
1 change: 1 addition & 0 deletions src/eva/assistant/pipecat_server.py
Original file line number Diff line number Diff line change
Expand Up @@ -368,6 +368,7 @@ async def _realtime_tool_handler(params) -> None:
alm_client=alm_client,
audio_collector=audio_llm_audio_collector,
output_dir=self.output_dir,
pre_tool_speech=self.pipeline_config.pre_tool_speech,
llm_streaming=self.pipeline_config.llm_streaming,
full_audio_context=self.pipeline_config.audio_llm_params.get("full_audio_context", False),
)
Expand Down
2 changes: 2 additions & 0 deletions src/eva/assistant/pipeline/audio_llm_processor.py
Original file line number Diff line number Diff line change
Expand Up @@ -190,6 +190,7 @@ def __init__(
alm_client: BaseALMClient,
audio_collector: AudioLLMUserAudioCollector,
output_dir: Path | None = None,
pre_tool_speech: str = "off",
llm_streaming: bool = False,
full_audio_context: bool = False,
**kwargs,
Expand All @@ -207,6 +208,7 @@ def __init__(
audit_log=audit_log,
alm_client=alm_client,
output_dir=output_dir,
pre_tool_speech=pre_tool_speech,
llm_streaming=llm_streaming,
full_audio_context=full_audio_context,
)
Expand Down
76 changes: 75 additions & 1 deletion src/eva/metrics/experience/turn_taking.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,13 +35,19 @@
user_interruption.mean_yield_ms,
user_interruption.mean_yield_score
(the latter two only when rate > 0)
Pre-tool speech: pretoolspeech_rate (lead-in tool-call groups / total tool-call groups;
computed from audit_log.json directly so it works uniformly across
cascade/S2S/audio-LLM; omitted when there are no tool calls or the
audit log is unavailable — see ``_compute_pre_tool_speech_groups``)

All reported sub-metrics are consistent with the main score: ``mean_overlap_score``,
``mean_count_score``, and ``mean_yield_score`` aggregate exactly the per-turn scores
that feed into ``turn_taking.score``.
"""

import json
import statistics
from pathlib import Path
from typing import Any

from eva.metrics.base import CodeMetric, MetricContext
Expand All @@ -59,7 +65,7 @@ class TurnTakingMetric(CodeMetric):
description = "Turn-taking evaluation based on per-turn latency and interruption behavior"
category = "experience"
pass_at_k_threshold = 0.8
version = "v0.1"
version = "v0.2"

# --- Latency curve (piecewise linear). 0 outside [LATENCY_HARD_EARLY_MS, LATENCY_HARD_LATE_MS]. ---
# Ramp up 0 → 1 from LATENCY_HARD_EARLY_MS to LATENCY_SWEET_SPOT_LOW_MS.
Expand Down Expand Up @@ -235,6 +241,64 @@ def _compute_yield_ms(context: MetricContext, turn_id: int) -> float | None:
agent_stopped = prev_a_segs[-1][1]
return max(0.0, agent_stopped - user_barge_in) * 1000

@staticmethod
def _compute_pre_tool_speech_groups(context: MetricContext) -> list[bool] | None:
"""Return one bool per contiguous run ("group") of tool_call/tool_response entries.

True when the assistant spoke (non-empty content) since the previous group ended (or
since the start of the conversation) — i.e. it gave a pre-tool-speech lead-in before this
group of tool calls. Consecutive tool calls with no intervening speech (e.g. two tools
invoked back-to-back off one lead-in) count as a single group, matching the
"one lead-in per batch" prompt instruction (``agent.pre_tool_speech`` in
configs/prompts/simulation.yaml) — a multi-tool turn with one lead-in isn't penalized for
the tools that didn't get their own.

Reads ``audit_log.json`` directly from ``context.output_dir`` instead of
``context.conversation_trace``. The audit log carries every event's true, original
timestamp (when the assistant actually generated/spoke the text), whereas for S2S
``conversation_trace`` assistant entries come from the user simulator's STT transcription
of the spoken audio — timestamped only after the audio finishes playing. That transcription
can sort *after* a subsequent tool_call in the timestamp-ordered trace even when the agent
actually spoke first (confirmed against a real S2S transcript), so ``conversation_trace``
order cannot be trusted for this signal on S2S. Reading the audit log directly sidesteps
that entirely and works uniformly across cascade, S2S, and audio-LLM.

Returns None when ``audit_log.json`` is missing, unreadable, or has no transcript — callers
should treat that as "unknown" rather than "no lead-ins".
"""
audit_log_path = Path(context.output_dir) / "audit_log.json"
try:
with open(audit_log_path) as f:
audit_log = json.load(f)
except (OSError, json.JSONDecodeError):
return None
transcript = audit_log.get("transcript")
if not transcript:
return None

groups: list[bool] = []
in_group = False
spoke_since_group = False
for entry in sorted(transcript, key=lambda e: int(e["timestamp"])):
message_type = entry.get("message_type")
if message_type in ("tool_call", "tool_response"):
if not in_group:
groups.append(spoke_since_group)
in_group = True
continue
# Ignore other event types (e.g. llm_call) without breaking the current group —
# only assistant/user speech should reset or extend the "spoke since group" state.
if message_type not in ("assistant", "user"):
continue
in_group = False
if message_type == "assistant":
content = entry.get("value", "")
if isinstance(content, str) and content.strip():
spoke_since_group = True
else:
spoke_since_group = False
return groups

@classmethod
def _per_turn_score_and_reason(
cls,
Expand Down Expand Up @@ -430,6 +494,16 @@ def _pct(p: float) -> float:
"user_interruption.mean_yield_score", round(statistics.mean(yield_scores), 4), True
)

# --- Pre-tool-speech lead-in rate ---
# Reads audit_log.json directly so it works uniformly across cascade, S2S, and audio-LLM
pre_tool_groups = cls._compute_pre_tool_speech_groups(context)
if pre_tool_groups:
sub["pretoolspeech_rate"] = _wrap(
"pretoolspeech_rate",
round(sum(pre_tool_groups) / len(pre_tool_groups), 4),
True,
)

# Token usage (from agent_perf_stats.csv)
mean_output_tokens = mean_agent_perf_stat(context.output_dir, "output_tokens")
mean_reasoning_tokens = mean_agent_perf_stat(context.output_dir, "reasoning_tokens")
Expand Down
17 changes: 11 additions & 6 deletions src/eva/models/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -197,6 +197,12 @@ class ModelConfig(BaseModel):
"off",
description="Prompt a model-generated lead-in before tool calls: 'off' or 'auto'.",
)

@field_validator("pre_tool_speech", mode="before")
@classmethod
def _normalize_pre_tool_speech(cls, value: str) -> str:
return value.lower() if isinstance(value, str) else value

llm_streaming: bool = Field(
False,
description="Stream Chat Completions output to TTS sentence-by-sentence.",
Expand Down Expand Up @@ -295,13 +301,12 @@ def _validate_latency_optimizations(self) -> "ModelConfig":
allowed = {"off", "auto"}
if self.pre_tool_speech not in allowed:
raise ValueError(f"pre_tool_speech must be one of {sorted(allowed)}, got '{self.pre_tool_speech}'")
# llm_streaming is honored by AUDIO_LLM too (via BaseALMClient.complete_stream); only
# pre_tool_speech / parallel_tool_calls remain CASCADE-only.
cascade_only_set = self.pre_tool_speech != "off" or self.parallel_tool_calls is not None
if cascade_only_set and self.pipeline_type != PipelineType.CASCADE:
# pre_tool_speech is honored by both CASCADE and AUDIO_LLM; llm_streaming by both
# (via BaseALMClient.complete_stream); only parallel_tool_calls remains CASCADE-only.
if self.parallel_tool_calls is not None and self.pipeline_type != PipelineType.CASCADE:
logger.warning(
"Cascade LLM flags (pre_tool_speech / parallel_tool_calls) apply only to the CASCADE "
f"pipeline; they will be ignored for pipeline_type={self.pipeline_type}."
"parallel_tool_calls applies only to the CASCADE pipeline; it will be ignored "
f"for pipeline_type={self.pipeline_type}."
)
return self

Expand Down
4 changes: 2 additions & 2 deletions tests/fixtures/metric_signatures.json
Original file line number Diff line number Diff line change
Expand Up @@ -92,8 +92,8 @@
"TurnTakingMetric": {
"name": "turn_taking",
"prompt_hash": null,
"source_hash": "fee8caa7adc7",
"version": "v0.1"
"source_hash": "1db130ae1aa6",
"version": "v0.2"
},
"UserBehavioralFidelityMetric": {
"name": "user_behavioral_fidelity",
Expand Down
Loading
Loading