Skip to content

[bot] Merge master/72655a20 into rel/dev - #1853

Merged
yenkins-admin merged 4 commits into
rel/devfrom
snapshot-master-72655a20-to-rel/dev
Oct 7, 2026
Merged

yenkins-admin merged 4 commits into
rel/devfrom
snapshot-master-72655a20-to-rel/dev

Conversation

@yenkins-admin

Copy link
Copy Markdown
Contributor

🚀 Automated PR to perform merge from master into rel/dev with changes up to 72655a2 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/37566819498).

myhoai and others added 4 commits October 6, 2026 17:59
When the answer carries no visualization part, the SSE client scores
the arguments of the last create_adhoc_visualization call. That path
was silent, so a turn that called the tool and answered without a
chart left no trace, although the user saw no chart either. It now
logs a warning with the number of calls.

Two fallback tests hand-wrote an `id` into the tool arguments, which
the real tool arguments never carry. That is why the missing-id crash
on this path passed CI. They now use the real shape, and the
synthesized-id test also checks the warning.

jira: QA-29242
risk: low
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tores it

A metric the eval deleted still showed up later in the same nightly.
On 2026-10-05, units_per_transaction was created by a passing
agent_conversations item on ecommerce_demo_sonnet55_anthropic: its
create_metric result said created_new: true, so the cleanup deleted
it, and it was still in the workspace after the job ended. Four other
combos showed the same thing that night.

create_metric in mcp-server reads the whole analytics model, adds the
metric and writes the model back. Tavern runs about fourteen workers
against one workspace, so a worker whose read came before our delete
puts the metric back when its write lands. The next item that asks
for that metric is told it already exists.

_delete_metric now checks after 3, 6 and 12 seconds that the metric
stayed deleted and deletes it again if it came back. A delete that
holds costs one check. Only a 404 counts as gone: a lookup that fails
otherwise deletes again rather than leave a restored metric behind, and
the fixed delays bound the retries. A failed delete is not rechecked. The
conversation evaluator uses the same function and gets the recheck
too.

This narrows the window rather than closing it. Moving the
metric-writing datasets to their own workspaces is in gdc-nas, and
the read-modify-write in create_metric is a product issue.

jira: QA-29117
risk: low
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The visualization evaluator compared the chart definition only. A
chart the agent built and never executed passed whenever the
definition was right, although the agent never saw the value: the
user got a number under a correct title that the reply did not state
(GDAI-2203, 22 of 24 Luna sessions).

Two signals, read from the tool calls the SSE stream already carries:

- executed: execute_visualization succeeded for a ref that a
  successful create_adhoc_visualization returned in the same
  conversation. It fails strict_pass only when the item sets
  expected_output.requires_execution. The gen-ai prompt tells the
  agent not to execute a chart the user only looks at, so a global
  gate would fail correct answers.
- stated_value_matches: whether a number in the reply is a rounding
  of a value the execution returned, or of its formatted string.
  None when the reply states no number, False when it states one
  and nothing ran. Reported only, because numbers in prose (years,
  "top 10") give false hits.

Both are scored in Langfuse on every item (assertion-vis-executed,
stated-value-matches). In detail they sit under "execution", so
quality_score, which counts every top-level boolean, changes only for
an item that gates on execution. The CLI reads the flag from the
item; evaluate_agentic_visualization takes it as requires_execution.

jira: QA-29248
risk: low
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(gooddata-eval): recheck metric deletes, report chart execution
@yenkins-admin
yenkins-admin merged commit be32dae into rel/dev Oct 7, 2026
1 check passed
@yenkins-admin
yenkins-admin deleted the snapshot-master-72655a20-to-rel/dev branch October 7, 2026 03:27
@coderabbitai

coderabbitai Bot commented Oct 7, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: b2da94c6-ffbe-4a74-88c2-ad5aa5f0c341

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 7, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.21488% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 84.17%. Comparing base (0a7b141) to head (72655a2).
⚠️ Report is 606 commits behind head on rel/dev.

Files with missing lines Patch % Lines
...al/src/gooddata_eval/core/agentic/visualization.py 50.00% 5 Missing ⚠️
...val/src/gooddata_eval/core/agentic/metric_skill.py 96.15% 1 Missing ⚠️
...src/gooddata_eval/core/evaluators/visualization.py 98.79% 1 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##           rel/dev    #1853      +/-   ##
===========================================
+ Coverage    84.12%   84.17%   +0.05%     
===========================================
  Files          333      333              
  Lines        23117    23235     +118     
===========================================
+ Hits         19447    19559     +112     
- Misses        3670     3676       +6     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants