Skip to content

Decorrelated the wait abort ISR test's interrupt source from the thread it samples - #743

Merged
fdesbiens merged 1 commit into
eclipse-threadx:devfrom
fdesbiens:fix/wait-abort-isr-decorrelation
Sep 16, 2026
Merged

fdesbiens merged 1 commit into
eclipse-threadx:devfrom
fdesbiens:fix/wait-abort-isr-decorrelation

Conversation

@fdesbiens

Copy link
Copy Markdown
Contributor

Follows #644 and #649, which bounded this test and lowered what it asks for. This
addresses the cause they worked around.

What the resonance actually is

The test waits for the periodic timer interrupt to land while _tx_thread_preempt_disable
is set. The same interrupt also puts the semaphore that wakes the thread being sampled,
so the two run in lockstep: the tick wakes thread 0, thread 0 does a fixed amount of work
and suspends, and the next tick arrives a fixed interval later with thread 0 in the same
place every time.

The window is a few instructions wide, so a sampler locked to the sampled thread's own
cycle can miss it indefinitely. Perturbing the handler's duration, which the handler
already does after 100 misses, shifts the phase but does not decorrelate the two.

The fix

_tx_linux_timer_interrupt waits on _tx_linux_timer_semaphore with a one-tick deadline
and delivers an interrupt early when the semaphore is posted — a mechanism the port
already relies on in _tx_thread_schedule. The test now runs a plain POSIX thread that
posts it, injecting interrupts at moments unrelated to the tick grid.

It posts only when nothing is outstanding, so interrupts can never be queued faster than
they are serviced and the injector cannot starve the system. It is confined to the Linux
simulation port by #ifdef __linux__, and no port file changes — the semaphore is a
global with external linkage, declared the same way tx_thread_schedule.c declares it.

Measured

Six build configurations of a branch carrying this change:

Windows reached Time
Before 1, 1, 2, 3, 3, 3 of 3 43–181 s each, three pinned at the budget
After 10 of 10, all six under 1 s each

The test job on that branch went from 20m48s to 4m27s.

The count goes back to ten on the same reasoning that lowered it to three — ask for what a
run can actually reach. The budget and the zero-window ceiling are untouched, so a run
that somehow still falls behind behaves exactly as it does today.

The SMP copy deliberately does not get this

#649 kept the two copies logically identical, and the budget and ceiling logic still is.
The injector is the exception, because only the tx copy has the problem it solves.

Measured on the SMP copy: 8.9 seconds with the injector against 0 to 1 without, for
the same twenty windows. It has never been in the slow mode, and the extra interrupt
traffic contends across the simulated cores. Applying it there would be a straight
regression, so its count stays at twenty and its file is unchanged.

Verification

  • tx suite 105/105, the test at 0.44 s; four direct runs reach 10 of 10 in 0–1 s.
  • SMP suite 117/117 unmodified, its own test 0–1 s over four runs.

…ad it samples

threadx_thread_wait_abort_and_isr_test waits for the periodic timer interrupt to
land while the preempt disable flag is set. The same interrupt also puts the
semaphore that wakes the thread being sampled, so the two run in lockstep: the tick
wakes thread 0, thread 0 does a fixed amount of work and suspends, and the next tick
arrives a fixed interval later with thread 0 in the same place every time. That is
the resonance the handler's own comment describes, and perturbing the handler's
duration only shifts the phase rather than breaking the correlation. The window
itself is a few instructions wide, so a sampler locked to the thread's own cycle can
miss it indefinitely, which is what eclipse-threadx#644 and eclipse-threadx#649 measured and worked around.

The simulator's timer thread waits on _tx_linux_timer_semaphore with a one-tick
deadline and delivers an interrupt early when the semaphore is posted, which the port
already relies on in _tx_thread_schedule. The test now runs a plain POSIX thread that
posts it, injecting interrupts at moments unrelated to the tick grid and sampling
thread 0 at arbitrary points in its cycle rather than the same one. The injector posts
only when nothing is outstanding, so interrupts can never be queued faster than they
are serviced. It is confined to the Linux simulation port and no port file changes.

The count of windows asked for goes back to ten, on the same reasoning that lowered
it to three: ask for what a run can actually reach. Measured over six build
configurations of a branch carrying this change, every one reaches ten of ten in under
a second, against one to three of three previously with three of the six pinned at the
180 second budget. The budget and the zero window ceiling are untouched, so a run that
somehow still falls behind behaves exactly as it does today.

The SMP copy deliberately does not get the injector, and its count stays at twenty. It
has never been in the slow mode, and injecting interrupts there measurably hurts:
8.9 seconds against 0 to 1 for the same twenty windows, which is the extra interrupt
traffic contending across the simulated cores. Only the tx copy has the problem this
solves, so only the tx copy changes; the budget and ceiling logic stays identical
between them.

Verified on this branch: the tx suite passes 105 of 105 with the test at 0.44 seconds,
and four direct runs reach ten of ten in 0 to 1 seconds. The SMP suite passes 117 of
117 unmodified, with its own test at 0 to 1 seconds over four runs.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
@fdesbiens
fdesbiens merged commit 24d1f27 into eclipse-threadx:dev Sep 16, 2026
10 checks passed
@fdesbiens
fdesbiens deleted the fix/wait-abort-isr-decorrelation branch September 16, 2026 13:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant