Repository navigation
[bot] Merge master/72655a20 into rel/dev - #1853
Conversation
When the answer carries no visualization part, the SSE client scores the arguments of the last create_adhoc_visualization call. That path was silent, so a turn that called the tool and answered without a chart left no trace, although the user saw no chart either. It now logs a warning with the number of calls. Two fallback tests hand-wrote an `id` into the tool arguments, which the real tool arguments never carry. That is why the missing-id crash on this path passed CI. They now use the real shape, and the synthesized-id test also checks the warning. jira: QA-29242 risk: low Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tores it A metric the eval deleted still showed up later in the same nightly. On 2026-10-05, units_per_transaction was created by a passing agent_conversations item on ecommerce_demo_sonnet55_anthropic: its create_metric result said created_new: true, so the cleanup deleted it, and it was still in the workspace after the job ended. Four other combos showed the same thing that night. create_metric in mcp-server reads the whole analytics model, adds the metric and writes the model back. Tavern runs about fourteen workers against one workspace, so a worker whose read came before our delete puts the metric back when its write lands. The next item that asks for that metric is told it already exists. _delete_metric now checks after 3, 6 and 12 seconds that the metric stayed deleted and deletes it again if it came back. A delete that holds costs one check. Only a 404 counts as gone: a lookup that fails otherwise deletes again rather than leave a restored metric behind, and the fixed delays bound the retries. A failed delete is not rechecked. The conversation evaluator uses the same function and gets the recheck too. This narrows the window rather than closing it. Moving the metric-writing datasets to their own workspaces is in gdc-nas, and the read-modify-write in create_metric is a product issue. jira: QA-29117 risk: low Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The visualization evaluator compared the chart definition only. A chart the agent built and never executed passed whenever the definition was right, although the agent never saw the value: the user got a number under a correct title that the reply did not state (GDAI-2203, 22 of 24 Luna sessions). Two signals, read from the tool calls the SSE stream already carries: - executed: execute_visualization succeeded for a ref that a successful create_adhoc_visualization returned in the same conversation. It fails strict_pass only when the item sets expected_output.requires_execution. The gen-ai prompt tells the agent not to execute a chart the user only looks at, so a global gate would fail correct answers. - stated_value_matches: whether a number in the reply is a rounding of a value the execution returned, or of its formatted string. None when the reply states no number, False when it states one and nothing ran. Reported only, because numbers in prose (years, "top 10") give false hits. Both are scored in Langfuse on every item (assertion-vis-executed, stated-value-matches). In detail they sit under "execution", so quality_score, which counts every top-level boolean, changes only for an item that gates on execution. The CLI reads the flag from the item; evaluate_agentic_visualization takes it as requires_execution. jira: QA-29248 risk: low Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(gooddata-eval): recheck metric deletes, report chart execution
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## rel/dev #1853 +/- ##
===========================================
+ Coverage 84.12% 84.17% +0.05%
===========================================
Files 333 333
Lines 23117 23235 +118
===========================================
+ Hits 19447 19559 +112
- Misses 3670 3676 +6 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🚀 Automated PR to perform merge from master into rel/dev with changes up to 72655a2 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/37566819498).