Skip to content

[Bugfix][Router] Keep avg_latency within the request-stats sliding window - #1110

Open
David-Wu1119 wants to merge 1 commit into
vllm-project:mainfrom
David-Wu1119:fix/avg-latency-window
Open

David-Wu1119 wants to merge 1 commit into
vllm-project:mainfrom
David-Wu1119:fix/avg-latency-window

Conversation

@David-Wu1119

Copy link
Copy Markdown

RequestStatsMonitor.get_request_stats reports per-engine statistics over the --request-stats-window sliding window. Before reading the QPS and TTFT monitors, it calls update_no_value(current_time) on each one so that samples older than the window are dropped. It never did that for the latency monitor, so avg_latency only shed old samples when a new request finished on that engine. Once an engine went idle, avg_latency stayed at the last value forever: the router kept exporting it (vllm:avg_latency) and logging it, while QPS and TTFT for the same engine had already dropped to 0 / -1.

For example, with a 10 s window, take one request that finishes 2 s after it started. 100 s later the stats read qps=0.0, ttft=-1, avg_latency=2.0. This change makes it avg_latency=-1, the same as TTFT, by dropping old latency samples the same way. (Open #1054 adds the same cleanup for avg_decoding_length; this is the latency counterpart.)

Tests:

  • New test_avg_latency_drops_requests_outside_the_window in src/tests/test_request_stats.py checks the example above. It fails on main (avg_latency is still 2.0) and passes with this change.
  • pytest src/tests: 241 passed, Python 3.12, pip install -e . plus pytest, pytest-asyncio and httpx.
  • pre-commit run on both files passes.

Found while auditing the router's request statistics with AI assistance; the fix and test were written with Claude Code.


  • Make sure the code changes pass the pre-commit checks.
  • Sign-off your commit by using -s when doing git commit
  • Try to classify PRs for easy understanding of the type of changes, such as [Bugfix], [Feat], and [CI].

🤖 Generated with Claude Code

…ndow

get_request_stats drops samples older than the sliding window from the QPS
and TTFT monitors before reading them, but not from the latency monitor,
so avg_latency averaged every request since the router started and stayed
frozen at the last value once an engine went idle. Drop old latency
samples the same way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: David-Wu1119 <133224895+David-Wu1119@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 06:21

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request ensures that requests outside the sliding window are correctly dropped when calculating average latency. This is achieved by calling update_no_value on the latency monitors within get_request_stats. Additionally, a new unit test test_avg_latency_drops_requests_outside_the_window has been added to verify this behavior. No review comments were provided, and the changes appear correct and complete.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants