Repository navigation
Conversation
Some models, e.g. Copilot's gpt-6-*, are rejected on /chat/completions
(`unsupported_api_for_model`) and only work through /responses.
- add a per model / per provider `api` option ("chat_completions" by
default, or "responses") for all OpenAI compatible providers; ignored
with a warning for anthropic, googleai and ollama
- resolution: `model.api` > `api` of the provider > Copilot lookup >
"chat_completions"
- Copilot: if nothing is set, ask GET /models once per session (lazily,
right after the bearer refresh) for each model's `supported_endpoints`
and use /responses for models that only support it. Failed lookups fall
back to /chat/completions and are retried on the next query
- Responses payload: system messages become `instructions`, the rest
`input`; temperature/top_p are not sent (reasoning models reject them);
`reasoning_effort`, `max_output_tokens` and `tools` are passed through
- endpoint: provider `responses_endpoint`, or `endpoint` with
/chat/completions replaced by /responses
- stream parser reads response.output_text.delta events
kai-leddy
force-pushed
the
feat/copilot-responses-api
branch
from
October 2, 2026 14:40
6f72a98 to
c4664c5
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note
Disclaimer: Both the code in this PR and the PR description were AI-assisted
Copilot's gpt-6 models fail on
/chat/completions:They only work on
/responses. So do a few other models in Copilot's list today, for examplegpt-5.5,gpt-5.6-luna,grok-4.6andmai-code-1.1-flash.What this does
apioption,"chat_completions"(the default) or"responses". Set it on an agent'smodel, or on a provider to change the default for all its models. It works for any OpenAI compatible provider. Anthropic, Google AI and Ollama ignore it and log a warning.apiisn't set, the plugin asksGET /modelsonce per session, right after the bearer refresh and only before the first query that needs it. It readssupported_endpointsfor each model. Models that only list/responsesare sent there. Models that list both stay on/chat/completions, so existing setups behave as before. If the lookup fails the plugin logs a warning, uses/chat/completionsand tries again on the next query.apiexplicitly skips the lookup.instructionsand the rest intoinput.temperatureandtop_pare not sent, since reasoning models reject them.reasoning_effort,max_output_tokensandtoolsare passed through if set on the model.responses_endpoint, orendpointwith/chat/completionsreplaced by/responses. If neither works, as with Azure's templated URL, the plugin logs a warning and sends nothing.response.output_text.deltaevents.Relation to #290
This overlaps with #290, which adds a separate
openai_respprovider. I went with an option on the model and provider instead. A provider entry also carries auth, and Copilot needs its own bearer refresh, so a second provider per vendor would copy that. Thetoolspassthrough comes from #290.Testing
The repo has no test suite, so I tested by hand.
apiresolution, plus a local fake SSE server for endpoint mapping, the/modelslookup, caching, concurrent first queries and a failing lookup./modelscall.gpt-6-astra,gpt-6-luna,gpt-6-sol,gpt-5.5,gpt-5.6-luna,gpt-5.4-miniandgrok-4.6went through/responses.gpt-5.4,claude-sonnet-5andgemini-3.8-flashstayed on/chat/completions. All returned a reply./responses.The README has a short note under the providers list. The vimdoc is left to the docgen action.