retryingOnlyBudgetOverruns
Runs scenario, retrying up to attempts times — but only when it fails by overrunning a runtime's wall-clock guest budget, and never on any other failure.
Why this exists (#1739)
WasmSandboxConfig.executionTimeout is a real wall-clock bound on real CPU-bound guest work. It cannot be driven by virtual time, so a test that exercises it is timing-gated by construction. That is correct for the runaway kernel the budget exists to catch. It is wrong for a well-behaved op that happens to share a runtime — and therefore the budget — with a runaway: reversing three bytes costs microseconds of guest work, but its deadline also covers being scheduled onto a CPU at all, and a saturated build host does not reliably manage that inside a few hundred milliseconds. Two of #1739's confirmed false reds were exactly this: an innocent invoke blowing a 250 ms / 200 ms budget at 1-minute load averages of 41, 52 and 71.6, and passing in isolation every time.
Raising the budget does not fix it. It buys a slower false red, and the runaway arm of the same test needs the budget tight so a non-conforming impl fails fast instead of burning the test host. Those two requirements are irreconcilable inside one WasmSandboxConfig — one runtime, one timeout — so the fix has to move to the assertion: the well-behaved arm asserts eventual success.
Why this does not weaken what the caller proves
Three independent filters, in the order they apply. Note that the type filter is the strongest and the message filter is narrower than it looks; the KDoc used to lead with a claim about trap text that is false for Wasm3WasmRuntime (its unreset-deadline trap is mapped to an overrun message, not a trap message), so the safety argument is stated here in dependency order instead.
Type. Only WasmExecutionException enters the
catchat all. A AssertionError — a wrong-bytesassertContentEquals, anassertFailsWiththat did not fire, anythingassertAllre-raises — and a us.tractat.kuilt.warp.WasmLoadException both propagate untouched. That covers every corrupted-state and missing-failure outcome without inspecting a single string.Message. Within WasmExecutionException, only BUDGET_OVERRUN_MARKER's full uniform phrasing is retried. A trap, an out-of-bounds ABI word, a rejected task on a dead worker and a stale JVM interrupt all read
"<phase> trapped: …"or"WASM kernel failed: …".Persistence. The filters above are not relied on to catch a runtime whose timeout no longer actually stops the guest, nor
Wasm3WasmRuntime's unreset deadline — both of those do present as overruns. They are caught because they are persistent: the worker stays busy or the deadline stays armed, so every attempt overruns and attempts of them are not enough.
What remains absorbed is therefore the transient overrun of a contended host, and nothing else. It used to also absorb #1802's residual-drain skew — a post-timeout op charged for the dying runaway's queue wait — but that is now fixed at the source: ChicoryWasmRuntime's worker proves itself free, or is discarded, before the timed-out caller returns.
The failure is self-describing
On exhaustion this reports every attempt's duration and then prices prepareReference — an equivalent well-behaved invocation on a fresh, generously-budgeted runtime — so the reader can separate contention from regression from the message alone, without re-running anything.
That price is deliberately not flattering to the host (#1810). It times the guest invocation alone, with construction and load hoisted out — parse and instantiate dominate a three-byte reverse, and the measurement it is compared against times an invoke on an already-loaded op — and it reports the fastest of REFERENCE_SAMPLES draws rather than one sample, since contention can only ever make a draw slower. Both corrections push the number down. That direction matters: an inflated reference reads as "at or near the budget ⇒ blame the host", the comfortable branch, when the honest reading may be "well under ⇒ REAL defect". Per the repo's debugging stance a contract-impossible value is a fork, and the measurement branch must not be the default.
An absorbed retry is not silent
Every overrun this absorbs is reported to emit as it happens, not accumulated for a report only the exhaustion path ever reads. Otherwise the drift from "no retry ever consumed" to "three of four consumed on every run" stays invisible until the day all four fail — at which point it reads as a fresh regression rather than as weeks of drift, which is exactly the failure #1739 exists to stop, one level up. It also gives the one thing #1801 left unproven a way to be measured: that review showed one attempt suffices at 1-minute load 74, not that four are enough there. Emitting absorbed overruns is what would ever tell us the tail rate had moved.
Parameters
What the scenario proves, for the failure message.
The WasmSandboxConfig.executionTimeout the scenario's runtime is configured with.
Maximum attempts; must be at least 1.
Where an absorbed overrun is reported, as it happens. Defaults to println, which Gradle captures into the per-test report — so a retry that eventually succeeds becomes visible in CI without failing anything. Override to capture it in a test.
The clock every duration in the report is measured against. Injected rather than read off the wall, so this function's own tests can pin what is timed without hoping a build machine is busy.
Builds an equivalent well-behaved guest invocation on a fresh, generously-budgeted runtime and returns the invocation alone. Construction and load happen inside this call and are therefore untimed; only the returned lambda is measured, so the reference prices the same thing the scenario's budget bounds.
The assertions to run. Must be safe to repeat.