(Disclaimer: had some help from Claude to put this together, but heavily edited)
Environment
- Elixir version (elixir -v): 1.19.5 (compiled with Erlang/OTP 28)
- Phoenix version (mix deps): 1.8.9
- Operating system: Linux (container), also reproduced on macOS
Actual behavior
The socket drainer is started inside the Endpoint's own supervisor, by Phoenix.Endpoint.Supervisor:
children =
config_children(mod, secret_conf, default_conf) ++
warmup_children(mod) ++
pubsub_children(mod, conf) ++
socket_children(mod, conf, :child_spec) ++
server_children(mod, conf, server?) ++
socket_children(mod, conf, :drainer_spec) ++
watcher_children(mod, conf, server?)
That means the drainer only starts running once the Endpoint itself begins to stop.
If your application supervision tree has any sibling child that your channels depend on, and that child is started after the Endpoint, the supervisor stops it first and the drainer has not yet touched a single socket.
The concrete case is Absinthe subscriptions. The documented setup is:
children = [
MyAppWeb.Endpoint,
{Absinthe.Subscription, MyAppWeb.Endpoint}
]
Absinthe.Subscription cannot be moved above the Endpoint, because Absinthe.Subscription.Proxy.init/1 calls pubsub.subscribe/1, which is MyAppWeb.Endpoint.subscribe/1, which calls pubsub_server!/0, which reads config(:pubsub_server) from the Endpoint's config ETS table.
That table is created by the Endpoint supervisor, so the Endpoint has to start first. Starting first means stopping last, so Absinthe.Subscription always dies before the drainer runs.
Every subscription message that arrives in that window crashes its channel:
** (RuntimeError) Pubsub not configured!
Subscriptions require a configured pubsub module.
(absinthe 1.11.0) lib/absinthe/phase/subscription/subscribe_self.ex:110: Absinthe.Phase.Subscription.SubscribeSelf.ensure_pubsub!/1
(absinthe_phoenix 2.0.3) lib/absinthe/phoenix/channel.ex:64: Absinthe.Phoenix.Channel.handle_in/3
(phoenix 1.8.9) lib/phoenix/channel/server.ex:332: Phoenix.Channel.Server.handle_info/2
Process Label: {Phoenix.Channel, Absinthe.Phoenix.Channel, "__absinthe__:control"}
This was filed in absinthe-graphql/absinthe_phoenix#46 since 2018. One of our pods logged 25 of these in the 8.5 seconds between SIGTERM and exit. We see zero while the pod is running normally.
The only way we found to get the right order is to call drainer_spec/1 ourselves and put it after the dependency:
# endpoint.ex - turn off the nested drainer
socket("/socket", MyAppWeb.UserSocket, websocket: true, drainer: false)
# application.ex - start it here instead, last, so it stops first
children = [
MyAppWeb.Endpoint,
{Absinthe.Subscription, MyAppWeb.Endpoint},
MyAppWeb.UserSocket.drainer_spec(endpoint: MyAppWeb.Endpoint, drainer: [shutdown: 10_000])
]
Supervisor.which_children/1 returns children in reverse start order, so this is the stop order on a running node after the change:
{:terminator, MyAppWeb.UserSocket} <-- drains sockets and channels FIRST
Absinthe.Subscription
MyAppWeb.Endpoint
This relies on things that are not public API:
drainer_spec/1 is @doc false (lib/phoenix/socket.ex). It is generated by use Phoenix.Socket for Phoenix.Endpoint.Supervisor to call.
drainer: false is undocumented. The socket/3 docs say :drainer takes "a keyword list or a custom MFA function returning a keyword list". false happens to work because the implementation reads it as if drainer = Keyword.get(opts, :drainer, []).
So the only working solution depends on two private behaviours, either of which could change in a patch release.
Expected behavior
Some supported way to control when the socket drainer runs relative to your own children. Some options:
- Make
drainer_spec/1 public and document drainer: false, so the snippet above is a supported pattern.
- Document, in the
socket/3 :drainer docs, that the drainer is started inside the Endpoint supervisor, and that anything your channels depend on must therefore be able to start before the Endpoint. Right now the docs describe the drainer's options and its behaviour relative to the Cowboy HTTP drainer, but never say where it sits.
- An option on the Endpoint, something like
drainer: [supervisor: :application], to opt into starting it outside.
Happy to send a PR for whichever of these you prefer.
(Disclaimer: had some help from Claude to put this together, but heavily edited)
Environment
Actual behavior
The socket drainer is started inside the Endpoint's own supervisor, by
Phoenix.Endpoint.Supervisor:That means the drainer only starts running once the Endpoint itself begins to stop.
If your application supervision tree has any sibling child that your channels depend on, and that child is started after the Endpoint, the supervisor stops it first and the drainer has not yet touched a single socket.
The concrete case is Absinthe subscriptions. The documented setup is:
Absinthe.Subscriptioncannot be moved above the Endpoint, becauseAbsinthe.Subscription.Proxy.init/1callspubsub.subscribe/1, which isMyAppWeb.Endpoint.subscribe/1, which callspubsub_server!/0, which readsconfig(:pubsub_server)from the Endpoint's config ETS table.That table is created by the Endpoint supervisor, so the Endpoint has to start first. Starting first means stopping last, so
Absinthe.Subscriptionalways dies before the drainer runs.Every subscription message that arrives in that window crashes its channel:
This was filed in absinthe-graphql/absinthe_phoenix#46 since 2018. One of our pods logged 25 of these in the 8.5 seconds between SIGTERM and exit. We see zero while the pod is running normally.
The only way we found to get the right order is to call
drainer_spec/1ourselves and put it after the dependency:Supervisor.which_children/1returns children in reverse start order, so this is the stop order on a running node after the change:This relies on things that are not public API:
drainer_spec/1is@doc false(lib/phoenix/socket.ex). It is generated byuse Phoenix.SocketforPhoenix.Endpoint.Supervisorto call.drainer: falseis undocumented. Thesocket/3docs say:drainertakes "a keyword list or a custom MFA function returning a keyword list".falsehappens to work because the implementation reads it asif drainer = Keyword.get(opts, :drainer, []).So the only working solution depends on two private behaviours, either of which could change in a patch release.
Expected behavior
Some supported way to control when the socket drainer runs relative to your own children. Some options:
drainer_spec/1public and documentdrainer: false, so the snippet above is a supported pattern.socket/3:drainerdocs, that the drainer is started inside the Endpoint supervisor, and that anything your channels depend on must therefore be able to start before the Endpoint. Right now the docs describe the drainer's options and its behaviour relative to the Cowboy HTTP drainer, but never say where it sits.drainer: [supervisor: :application], to opt into starting it outside.Happy to send a PR for whichever of these you prefer.