Repository navigation
Asked the wait abort test for three windows, and failed a run that reached none - #649
Merged
fdesbiens merged 1 commit intoAug 20, 2026
Conversation
…ached none
Four CI runs of the same tree, twenty configuration-runs in total, show this
test's budget being reached far more often than the first green run suggested,
and a pass being reported every time it was:
trace_build 3 of 10 windows in 121 seconds
disable_notify 3 of 10 windows in 121 seconds
default_coverage 4 of 10 windows in 121 seconds
stack_checking 7 of 10 windows in 121 seconds
trace_build 0 of 10 windows in 121 seconds
disable_notify 7 of 10 windows in 121 seconds
stack_checking 3 of 10 windows in 121 seconds
Seven of twenty, and the shortfall message only ever reaches an artifact:
ctest is run with --output-on-failure, so a passing test's output is not in the
job log at all. The suite has been quietly losing most of this test's coverage
in whole configurations and reporting green.
The loop runs in two modes, not one. A window arrives in milliseconds in the
fast mode, and costs between 17 and 40 seconds in the slow one, with nothing in
between across those twenty runs. Ten windows are therefore unreachable inside
any budget worth having: at 40 seconds each that is 400 seconds, and the
unbounded runs measured before any of this took up to 726. Raising the budget
to cover the slow mode would trade a quiet loss of coverage for five
configurations approaching the sixty minute step timeout.
So ask for what a run can reach. Three windows cost 51 to 120 seconds in the
slow mode and under a second in the fast one, and the later hits repeat what
the first ones establish, so what is given up is small. The budget goes to 180
seconds because three windows at the worst rate measured is exactly the 120 it
was, which would have truncated at two.
The count is printed on every run rather than only on a short one. A number
that appears only on shortfall cannot be told apart from a number nobody
recorded.
Reaching the window no times at all is a different matter, and was the worst of
the seven. The check after the loop compares semaphore bookkeeping that a
window has to have touched to mean anything, so a run that reached none of them
compares a counter against the value it was initialised to and reports a pass
having verified nothing. That run now keeps trying to a 300 second ceiling, and
fails if it still has not reached the window. A genuine resonance that holds
for five minutes is worth a failure; the old behaviour was worth nothing.
The SMP copy keeps its count of twenty. It reaches them in under half a second
in all five of its configurations, in all four runs, so the slow mode has never
been observed there and the coverage is free. Both copies get the ceiling and
the unconditional report, so the logic stays identical between them.
Verified locally on all five configurations: the test reaches 3 of 3 in 5 to 14
seconds, and the full suites pass 96 of 96 and 110 of 110 run one test at a
time. With the handler's window made unreachable and the ceiling lowered to 5
seconds, the test stops after 6 seconds, prints the count it reached, and
reports ERROR eclipse-threadx#8 with the harness recording a failure rather than a pass. With
the count raised past what the budget allows, a run that reaches two windows
still passes, so falling short and reaching nothing stay distinct. The
TX_NOT_INTERRUPTABLE branch, which no configuration in either suite builds, was
compile-checked in both copies with the configurations' own compile commands.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
fdesbiens
added a commit
that referenced
this pull request
Sep 16, 2026
…ad it samples (#743) threadx_thread_wait_abort_and_isr_test waits for the periodic timer interrupt to land while the preempt disable flag is set. The same interrupt also puts the semaphore that wakes the thread being sampled, so the two run in lockstep: the tick wakes thread 0, thread 0 does a fixed amount of work and suspends, and the next tick arrives a fixed interval later with thread 0 in the same place every time. That is the resonance the handler's own comment describes, and perturbing the handler's duration only shifts the phase rather than breaking the correlation. The window itself is a few instructions wide, so a sampler locked to the thread's own cycle can miss it indefinitely, which is what #644 and #649 measured and worked around. The simulator's timer thread waits on _tx_linux_timer_semaphore with a one-tick deadline and delivers an interrupt early when the semaphore is posted, which the port already relies on in _tx_thread_schedule. The test now runs a plain POSIX thread that posts it, injecting interrupts at moments unrelated to the tick grid and sampling thread 0 at arbitrary points in its cycle rather than the same one. The injector posts only when nothing is outstanding, so interrupts can never be queued faster than they are serviced. It is confined to the Linux simulation port and no port file changes. The count of windows asked for goes back to ten, on the same reasoning that lowered it to three: ask for what a run can actually reach. Measured over six build configurations of a branch carrying this change, every one reaches ten of ten in under a second, against one to three of three previously with three of the six pinned at the 180 second budget. The budget and the zero window ceiling are untouched, so a run that somehow still falls behind behaves exactly as it does today. The SMP copy deliberately does not get the injector, and its count stays at twenty. It has never been in the slow mode, and injecting interrupts there measurably hurts: 8.9 seconds against 0 to 1 for the same twenty windows, which is the extra interrupt traffic contending across the simulated cores. Only the tx copy has the problem this solves, so only the tx copy changes; the budget and ceiling logic stays identical between them. Verified on this branch: the tx suite passes 105 of 105 with the test at 0.44 seconds, and four direct runs reach ten of ten in 0 to 1 seconds. The SMP suite passes 117 of 117 unmodified, with its own test at 0 to 1 seconds over four runs. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
fdesbiens
added a commit
that referenced
this pull request
Sep 18, 2026
threadx_thread_wait_abort_and_isr_test's ThreadX SMP copy still waits for a probabilistic window with no bound, no failure path and no record of what it reached. A run that settles into the resonance the ISR comment describes never ends and is killed by ctest at its timeout with nothing usable in the log. The ThreadX copy was given a bound in #644 and #649; only that copy was changed, and the SMP half has been waiting since. This is those two commits' SMP half, applied unchanged. The budget is wall clock rather than ticks, because a tick arrives only when the port's timer thread runs and so falls behind real time under exactly the load that makes the wait long. The count asked for stays at twenty: this copy reaches twenty windows in under half a second in all eight configurations, so there is nothing to gain by asking for fewer. A run reaching the window no times at all now fails rather than passing on bookkeeping no window ever touched, and every run prints the count it reached. Verified in default_build_coverage: a normal run reaches 20 of 20 and passes; with the count raised past reach and the budget cut to three seconds it stops at four seconds, reports 114 of 4000000 and passes; with the window made unreachable it reports 0 of 20, prints ERROR #8 and exits 1. The test passes in each of the eight build configurations. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
threadx_thread_wait_abort_and_isr_testwas given a wall clock budget in #644 so an unreachable race window could not hang a run. Four runs of the same tree since then show the budget being reached far more often than the first green run suggested, and a pass being reported every time it was.The measurements
Twenty configuration-runs of the non-SMP copy. Seven ran out of budget:
Every one of those reported a pass, and the shortfall message reaches only the uploaded artifact — ctest runs with
--output-on-failure, so a passing test's output is not in the job log at all. The suite has been losing most of this test's coverage in whole configurations, in three of them at once, and reporting green.Ten windows are not reachable inside any sensible budget
The loop runs in two modes rather than one, with nothing in between across those twenty runs:
Ten windows in the slow mode is 400 seconds at the worst rate, and the unbounded runs measured before #644 reached 726. Raising the budget far enough to cover that would put five configurations within reach of the sixty minute step timeout added in #641 — trading a quiet loss of coverage for a loud timeout risk.
The change
Ask for what a run can reach.
WAIT_ABORT_WINDOWS_WANTEDgoes from 10 to 3 in the non-SMP copy. Three windows cost 51 to 120 seconds in the slow mode and under a second in the fast one. The rationale committed with #644 applies directly: the value is in reaching the window at all, and the later hits repeat what the first ones establish.Budget 120 → 180 seconds. Three windows at the worst rate measured is exactly the 120 it was, which would have truncated at two.
Report the count on every run, not only on a short one. A number that appears only on shortfall cannot be told apart from a number nobody recorded.
Reaching no window at all now fails. That case is different in kind from falling short: the check after the loop compares semaphore bookkeeping that a window has to have touched to mean anything, so a run that reached none compares a counter against the value it was initialised to and reports a pass having verified nothing. Such a run keeps trying to a 300 second ceiling and fails with
ERROR #8if it still has not reached the window. A resonance that holds for five minutes is worth a failure.The SMP copy keeps its count of twenty. It reaches them in under half a second in all five of its configurations across all four runs, so the slow mode has never been observed there and that coverage is free. Both copies get the ceiling and the unconditional report, so the logic stays identical between them — the divergence in #648 was a reminder of what happens when it does not.
Verification
In CI, in the slow mode. The verification run happened to land in the slow mode, which is the case that matters, and every configuration reached full coverage inside the new budget:
All five SMP configurations reached 20 of 20 in 0 to 1 second. Every suite passed in full: ThreadX 96 of 96 and SMP 110 of 110 across five configurations each, FreeRTOS 3 of 3, with suite totals of 36 to 70 seconds per configuration.
The new failure path fires, and only when it should. With the handler's window made unreachable and the ceiling lowered to 5 seconds, the test stops after 6 seconds, prints the count it reached, and reports
ERROR #8with the harness recording a failure rather than a pass. With the count raised past what the budget allows, a run that reaches two windows still passes — so falling short and reaching nothing stay distinct.Locally, all five configurations reach 3 of 3 in 5 to 14 seconds, and both suites pass 96 of 96 and 110 of 110 run one test at a time.
The
TX_NOT_INTERRUPTABLEbranch is compile-checked in both copies using the configurations' own compile commands. No configuration in either suite builds it, so it is not covered by a normal run;condition_countis expected to be zero there, which is why the new failure is compiled out of that path.Known, and deliberately not addressed here
threadx_thread_delayed_suspension_testhas the same shape of hole. Since #645 it skips its dependent check when its window is not reached, so a zero-window run there is also a pass that verified nothing. It has never fallen short in twenty configuration-runs — 4.06 seconds against a 120 second budget at worst — so there is no measured problem to fix, and it is left alone rather than changed on the strength of an argument by analogy.One unrelated intermittent turned up while measuring.
threadx_smp_random_resume_suspend_exclusion_pt_testfailed withERROR #7on a first attempt intrace_buildand passed on retry, on a sixteen core machine running the suite one test at a time. It did not recur in any CI run, including one with--repeat until-pass:1, so it is recorded here rather than diagnosed.