Skip to content

feat: support the OpenAI Responses API, fixes Copilot gpt-6 models - #302

Open
kai-leddy wants to merge 1 commit into
Robitx:mainfrom
kai-leddy:feat/copilot-responses-api
Open

kai-leddy wants to merge 1 commit into
Robitx:mainfrom
kai-leddy:feat/copilot-responses-api

Conversation

@kai-leddy

@kai-leddy kai-leddy commented Oct 2, 2026 •

Copy link
Copy Markdown

Note

Disclaimer: Both the code in this PR and the PR description were AI-assisted

Copilot's gpt-6 models fail on /chat/completions:

400: {"message":"model \"gpt-6-astra\" is not accessible via the /chat/completions endpoint","code":"unsupported_api_for_model"}

They only work on /responses. So do a few other models in Copilot's list today, for example gpt-5.5, gpt-5.6-luna, grok-4.6 and mai-code-1.1-flash.

What this does

  • Adds an api option, "chat_completions" (the default) or "responses". Set it on an agent's model, or on a provider to change the default for all its models. It works for any OpenAI compatible provider. Anthropic, Google AI and Ollama ignore it and log a warning.
  • For Copilot, when api isn't set, the plugin asks GET /models once per session, right after the bearer refresh and only before the first query that needs it. It reads supported_endpoints for each model. Models that only list /responses are sent there. Models that list both stay on /chat/completions, so existing setups behave as before. If the lookup fails the plugin logs a warning, uses /chat/completions and tries again on the next query.
  • I chose a lookup over a hard coded model list because a list would go stale as Copilot adds models. Setting api explicitly skips the lookup.
  • The Responses payload turns system messages into instructions and the rest into input. temperature and top_p are not sent, since reasoning models reject them. reasoning_effort, max_output_tokens and tools are passed through if set on the model.
  • The endpoint is the provider's responses_endpoint, or endpoint with /chat/completions replaced by /responses. If neither works, as with Azure's templated URL, the plugin logs a warning and sends nothing.
  • The stream parser reads response.output_text.delta events.

Relation to #290

This overlaps with #290, which adds a separate openai_resp provider. I went with an option on the model and provider instead. A provider entry also carries auth, and Copilot needs its own bearer refresh, so a second provider per vendor would copy that. The tools passthrough comes from #290.

Testing

The repo has no test suite, so I tested by hand.

  • Unit checks of payload building and api resolution, plus a local fake SSE server for endpoint mapping, the /models lookup, caching, concurrent first queries and a failing lookup.
  • Live Copilot, ten models in one cold start run with a single /models call. gpt-6-astra, gpt-6-luna, gpt-6-sol, gpt-5.5, gpt-5.6-luna, gpt-5.4-mini and grok-4.6 went through /responses. gpt-5.4, claude-sonnet-5 and gemini-3.8-flash stayed on /chat/completions. All returned a reply.
  • Not tested against Azure or against OpenAI's own /responses.

The README has a short note under the providers list. The vimdoc is left to the docgen action.

Some models, e.g. Copilot's gpt-6-*, are rejected on /chat/completions
(`unsupported_api_for_model`) and only work through /responses.

- add a per model / per provider `api` option ("chat_completions" by
  default, or "responses") for all OpenAI compatible providers; ignored
  with a warning for anthropic, googleai and ollama
- resolution: `model.api` > `api` of the provider > Copilot lookup >
  "chat_completions"
- Copilot: if nothing is set, ask GET /models once per session (lazily,
  right after the bearer refresh) for each model's `supported_endpoints`
  and use /responses for models that only support it. Failed lookups fall
  back to /chat/completions and are retried on the next query
- Responses payload: system messages become `instructions`, the rest
  `input`; temperature/top_p are not sent (reasoning models reject them);
  `reasoning_effort`, `max_output_tokens` and `tools` are passed through
- endpoint: provider `responses_endpoint`, or `endpoint` with
  /chat/completions replaced by /responses
- stream parser reads response.output_text.delta events
@kai-leddy
kai-leddy force-pushed the feat/copilot-responses-api branch from 6f72a98 to c4664c5 Compare October 2, 2026 14:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant