Conversation
Readiness down-detection relies on the Kafka client's hasReadyNodes(), which only flips after connection-failure detection bounded by request.timeout.ms (default 30s). The 20s Awaitility window was shorter than that bound, so a killed-broker connection was not always detected in time, causing intermittent ConditionTimeout failures in CI (all 3 rerun attempts failed). Widen the await window to 45s and add a method-level @timeout(60) (overriding the class-level @timeout(30)) so detection can complete reliably. Both are upper bounds and Awaitility returns as soon as the condition is met, so passing runs are not slowed. Also drop unnecessary public modifiers per JUnit 5 conventions. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Claus Ibsen <claus.ibsen@gmail.com>
|
🌟 Thank you for your contribution to the Apache Camel project! 🌟 🐫 Apache Camel Committers, please review the following items:
|
|
🧪 CI tested the following changed modules:
🔬 Scalpel shadow comparison — Scalpel: 1 tested, 0 compile-only — current: 10 all testedMaveniverse Scalpel detected 1 affected modules (current approach: 10). Modules only in current approach (9)
Skip-tests mode would test 1 modules (1 direct + 0 downstream), skip tests for 0 (generated code, meta-modules) Modules Scalpel would test (1)
All tested modules (9 modules)
|
…teFixture reflection works on Java 25 The class was made package-private as a JUnit 5 convention cleanup, but the test-infra CamelContextExtension invokes the @RouteFixture createRouteBuilder method reflectively. On Java 25, Method.invoke on a package-private receiver class throws IllegalAccessException even for a public method, breaking all tests at fixture setup. Restore the public modifier (the three sibling HealthCheck ITs are already public for the same reason). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Claus Ibsen <claus.ibsen@gmail.com>
|
provided #25645 to fix the test which is not flaky |
|
Closing this in favor of #25645, which fixes the actual root cause. The timeout increase here can't work: the failure is deterministic, not a timing flake. This was introduced by CAMEL-24387 (#22294), which migrated these health-check ITs to the singleton service. #25645 correctly reverts just this base class back to Verified locally with a real Kafka container:
Also filing a follow-up: under Claude Code on behalf of davsclaus |
Issue
CAMEL-24466
KafkaConsumerHealthCheckIT.testReadinessWhenDownfails intermittently in CI withConditionTimeout— the readiness health check does not reportDOWNwithin the 20s Awaitility window after the broker is shut down. In the reported run it failed all 3 attempts (initial + 2rerunFailingTestsCountreruns), so it was not masked as flaky.Root cause
Readiness down-detection is driven by
KafkaFetchRecords.isReady(), which (whileconnectedstaystrue) relies on the Kafka client'sConsumerNetworkClient.hasReadyNodes(now). When a live broker is killed without a prompt TCP reset (typical for a container stop under CI load), the client only marks the node not-ready once the in-flight request times out — bounded by the consumer'srequest.timeout.ms, which defaults to 30s (KafkaConfiguration.consumerRequestTimeoutMs = 30000).The test's 20s await window is shorter than that 30s bound, so detection sometimes has not completed when Awaitility gives up → flaky
ConditionTimeout.This is corroborated by the sibling tests that are not flaky (
KafkaConsumerBadPortHealthCheckIT,KafkaConsumerUnresolvableHealthCheckIT): they connect to a bad/unresolvable endpoint, so a connection never becomes ready andhasReadyNodes()returnsfalsealmost immediately — 20s is ample there. OnlytestReadinessWhenDownestablishes a READY connection first and then kills the broker, exposing the ~30s detection delay.Fix
request.timeout.msdetection bound.@Timeout(60)(overriding the class-level@Timeout(30), which would otherwise kill the method first).Both are upper bounds, not sleeps: Awaitility returns as soon as the condition is met, and JUnit only fails past the ceiling, so passing runs are not slowed. Eliminating the fail-then-rerun cycle tends to reduce total CI time for this class.
Also dropped the unnecessary
publicmodifiers on the touched class/method per JUnit 5 conventions.Testing
mvn -DskipTests installoncamel-kafkapasses (test sources compile;formatter:format/impsort:sortproduce no changes).generate-postcompilereported everything up to date).Claude Code on behalf of davsclaus