Topic

Openrouter

4 episodes

  1. Ep 903

    Z.ai launches GLM 5.3 Flash under MIT license

    GLM-5.3-Flash drops today under MIT license — 320 billion parameters, 18 billion active, one million token context, and it was hiding in plain sight as Ox Alpha on OpenRouter all week. Edmund and Geffen dig into the architecture, the benchmark claims, the GLM-5.3 weights bet that's now two days from settling, and whether a model that costs fifteen cents per million input tokens actually changes the open-weight story.

  2. Ep 889

    Stripe Payments Openrouter Singularity

    Stripe says January 1, 2026 marked the beginning of a major technological and economic inflection point, using its business data as evidence. The more practical move may be its acquisition of OpenRouter, connecting model routing to the payments and control infrastructure Stripe already owns.

  3. Ep 812

    Qwen 3.7 Flash review: a $0.03 vision model with a catch

    Cathy leads a skeptical take on Qwen 3.7 Flash — the $0.03 vision model from Alibaba that looks like a pricing breakthrough until you read the fine print. The tiered pricing structure, near-zero independent benchmarks, a ninety-second P99 latency tail, and an eight-point-nine percent tool error rate make it a much narrower product than the headline suggests. Jessica steelmans the volume-processing use case and the genuine competitive pressure it puts on the cheap tier, but neither host pretends the transparency gap isn't a real problem.

  4. Ep 628

    How to Run Open Source AI Models

    Sid Saladi argues that frontier AI vendors (Claude, GPT) bundle model, compute, access, and application into one proprietary stack—trapping users in unpredictable pricing and competitive capture. The counter: open-weight models like GLM-5.2, DeepSeek V4, Qwen, and Kimi are now frontier-adjacent in capability (GLM-5.2 beats GPT-5.5 on coding benchmarks, matches Opus 4.8 on others) and cost roughly one-sixth as much. The real problem isn't model quality anymore; it's that companies like Tesla, Uber, and Meta are hemorrhaging money on metered AI because they can't decouple the stack. The guide walks four layers—model, compute, access, harness—and shows how to own each one deliberately instead of letting a vendor own all four by default.