I Built a Safety Filter With No Off Switch, and It Blocked Every Closed Model

The requirement sounded simple: never send a request to a provider serving a quantized copy of a model. I implemented it as a hard invariant, and in doing so I made every closed-weight model in our catalog impossible to use. Every run against those models came back as a 400. Here’s what went wrong and what I should have done.

What the code actually did

The provider candidate filter runs at enqueue time. Before it evaluates anything else, it tries to determine the model’s native quantization level. There are exactly two ways it can learn that: a hand-set catalog value, or the best precision declared by any upstream endpoint. If neither produces an answer, the filter refuses the whole enqueue.

For open-weight models this works fine. Somebody, somewhere, publishes a precision. For closed models it falls apart completely. Every endpoint reports its quantization as unknown, and the hand-set catalog field is validated against a fixed list of concrete levels, so “unknown” isn’t a value I could type in either. The filter therefore failed at step one with a NoNativeLevel error.

The cruel part is that there was an escape hatch. The policy struct has an allow list for providers with unknown quantization, and I had populated it correctly. But that check sits two steps later in the pipeline. The native-level determination runs first and refuses before the allow list is ever consulted. The configuration looked right. The code path was unreachable.

Why I built it that way

I took the requirement literally and turned it into a mandatory invariant: every candidate must be provably at the native level. That framing carried an assumption I never examined, which is that some provider always publishes a precision. I treated unknown as a per-provider anomaly to be steered around, not as a state an entire model can be in.

That assumption shaped the whole policy struct. Every knob I gave it narrows the candidate list: native level, price ceiling, provider bans, the per-provider unknown allow list. There is no knob that turns the filter off. Then I validated the hand-set native level strictly against the enumerated levels, which closed the one remaining way out.

My tests covered the open-weight shapes, because those were the shapes I had in my head while writing the filter. A curated model with undisclosed weights was never exercised end to end.

What I should have done differently

Give a safety filter an explicit off switch on day one. A filter that can only be tightened will eventually refuse something legitimate. The off switch is not a compromise of the safety goal, it’s the acknowledgement that the operator knows things the filter doesn’t. If I had built the toggle first and the invariant second, this bug could not have existed.

When a filter can refuse, enumerate the inputs that must still pass. I spent my design effort on what should be blocked and none on what must survive. “Closed model, every endpoint reports unknown” was the obvious case and I missed it because I never wrote the list down.

Validate escape hatches against the failure they’re meant to escape. I tested the unknown-quantization allow list for the case of one unknown provider among known ones. I never tested it for the case where every provider is unknown, which is the only case where it actually matters. A test that exercises a hatch under mild conditions tells you nothing about whether it opens under load.

Treat order-dependent filters as a place dead code hides. Reading the allow list in isolation, it looked correct. It was correct. It just could never fire for the models that needed it, because an earlier stage refused first. Any pipeline where stage N can terminate before stage N+1 needs at least one test that proves stage N+1 is reachable for the inputs it was written for.

A good error message is not a substitute for a working design. The refusal explained its own mechanism clearly. It told the operator exactly which value could not be determined and exactly which field could not be set to fix it. I was briefly proud of that message. In retrospect it was the bug report: a message that precisely describes an unconfigurable dead end is telling you the design has no answer, and I should have read it that way the first time I saw it.

The fix

I’m adding a catalog-level toggle that disables the quantization filter for a given model. With it off, the native-level determination step is skipped entirely and every endpoint clears the quantization gate. Everything else still applies, including the price ceiling, cache-read pricing, parameter support checks, and the provider ban list. Turning off the quantization filter turns off the quantization filter, nothing more.

The change touches the core filter, the backend policy plus a migration, the model configuration page, the candidates endpoint, and the docs. I’m also adding the test I should have written originally: a model where every endpoint reports unknown, asserted to enqueue successfully with the toggle off and to refuse with it on.

The broader habit I’m changing

The pattern here isn’t really about quantization. It’s that I converted a product requirement phrased as “never” into a code-level invariant with no override, and then validated my way into a corner. “Never do X” from a stakeholder almost always means “don’t do X by default, and make it deliberate when we do.” Those are different programs. I wrote the wrong one, and the strictness I was proud of turned into an outage.

Going forward, when I implement a rule that can reject work, I’m writing down three things before the implementation: what must be rejected, what must still pass, and how an operator overrides the rule when I’ve gotten the first two wrong. That third item is the one I skipped, and it’s the one that would have saved me.