Live demo

5,000 tool calls in. 788 out.

5,000 calls to search_orders under a 30s timeout. 788 of them were still running when the config gave up. Kaplan-Meier treats those as censored, not as slow, and reads off the second at which another second of waiting stops paying.

Timeout in the config
30s
Where waiting stops paying
6s
next-second completion under 2.0%
Completions cut
0.4%
of calls that finish under the current timeout
Saved per hung call
24s
80.0% less wait on hangs
Given the call has already waited t seconds
WaitedStill runningCalls at riskCompletes in next 1s95% upperVerdict
1s68.2%3,44152.4%54.1%
wait
2s32.5%1,62734.5%36.8%
wait
3s21.3%1,06618.0%20.4%
wait
4s17.4%8746.1%7.9%
wait
5s16.4%8192.0%3.1%
wait
6s16.1%8030.9%1.8%
stop
7s15.9%7960.9%1.8%
stop
8s15.8%7890.0%0.5%
stop
10s15.8%7890.1%0.7%
stop
15s15.8%7880.0%0.5%
stop
20s15.8%7880.0%0.5%
stop
Wait per hung call today
30s
Wait per hung call
6s

788 hung calls in this window paid 23,640s of wait. At 6s they pay 4,728s, and 0.4% of completing calls get cut to buy that.

The p99 rulep99 of completed calls is 4.76s

The other number people reach for. It is a quantile of the calls that came back, so it cuts 1.0% of good calls by construction, whether or not anything is hanging, and it says nothing about the 15.8% that never return. Kaplan-Meier cuts 0.4% here because it can see that the curve has gone flat.

Diagnostic1 completions between 10s and 30s

The method assumes the calls that hit the timeout were hangs, not slow completions. If they were slow completions, the curve would keep sloping between ten seconds and the timeout instead of going flat at 15.8%. It went flat.