source avatarPetrAnto

Share

OpenRouter Jev Router: The Cache Is the Real Routing Problem OpenRouter says its new Jev Router solved 237 of 423 agent tasks versus 130 for Auto. One early live test found the same routing logic could keep paying premium model prices to answer “what color is a banana?” ---------------------- OpenRouter launched `typesafe/jev-router` yesterday. Jev does not answer your prompt. It is a tiny decision model that looks at the conversation, scores things like task difficulty, required precision and expected benefit from a larger model, then decides which LLM and reasoning effort should handle the next turn. OpenRouter says that architecture solved 82% more tasks than its existing Auto router across four agent benchmarks. That would be a huge improvement. There is just one problem: the benchmark names, cost per solved task and full comparison data are not public yet. And the most interesting early criticism is about something Jev was explicitly designed to optimize 👇 🔹The Cache Is the Product🔹 Model routing sounds simple: - Easy question → cheap model - Hard question → expensive model Agent sessions make that much harder. Changing models can destroy the prompt cache, forcing the new model to process the entire conversation again. Jev Router therefore asks a smarter question: Is the expected gain from switching worth the cost of losing the cache? It can also change reasoning effort while staying on the same model. That is genuinely interesting product engineering. But cache economics work in both directions. 🔹The Banana Problem🔹 Zach Moskow tested a four turn conversation. The router started cheaply for “capital of France,” moved to a premium model for a number theory proof... Then stayed there for: “What’s 7×8?” And: “What color is a banana?” His banana turn cost roughly 140 times the first turn. Why? Downgrading also sacrifices the cache. So the same rule designed to prevent wasteful model switching can create a one way premium trap. The router becomes sticky upward. A smart routing policy can optimize every individual switch and still produce a bad session level bill. 🔹Jev Itself Is Not The Problem🔹 Jev looks genuinely useful for what it was built to do. On OpenRouter's Banking77 classification test, Jev reached 81.0% accuracy versus Claude Opus 5 at 84.4%, while running around 13 times faster and costing roughly 22 times less. That is compelling for typed decisions. But classification performance does not automatically prove that prompt text contains enough information to judge real task difficulty. “Fix this bug” could mean two lines. Or two days. A router that cannot see the repository, runtime state, tools or future branches of the task is estimating complexity from an incomplete picture. That distinction matters much more for agents than for support ticket classification. 🔹Auto Versus Jev Is Really Two Philosophies🔹 OpenRouter Auto behaves more like a market index. It classifies the task, then uses recent OpenRouter spending patterns to choose models inside a cost band, with fallbacks when routing fails. Jev Router makes a more explicit judgment about difficulty, precision, reasoning effort and cache economics. Auto asks: what models are people successfully paying for? Jev asks: how hard does this look, and is switching worth it? The second approach is intellectually cleaner. Whether it is economically better still needs to be proven. 🔹What I Would Actually Do🔹 ▫️ Test complete workflows rather than isolated prompts. 📝 Compare solve rate, total session cost, model switches, cache hits and time to first token across the same real agent tasks. ▫️ Specifically test what happens after the hardest turn. 📝 Give the router an expensive reasoning task followed by several trivial requests and watch whether model cost steps back down. ▫️ Do not confuse free routing with cheap inference. 📝 Jev Router itself costs nothing to invoke, but the downstream model it selects is still the bill that matters. ▫️ Keep predictable workloads explicit for now. 📝 If one model already reliably handles a workflow inside your budget, routing adds another policy layer that must earn its complexity. 🔹What Would Convince Me🔹 I want a named benchmark with the same agent harness comparing: - solve rate - total cost - cost per solved task - cache hit rate - model switches - TTFT - failure rate The 237 versus 130 result is intriguing. Without the economics beside it, it is still only half the benchmark. The breakthrough is not choosing a smarter model. It is knowing when staying on it costs more than switching away.

No.0 picture
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.