Domain

New Models

95 episodes

  1. Ep 786

    Introducing Claude Opus 5

    Anthropic ships Claude Opus 5 — a model that hits near-Fable-5 performance on coding and knowledge work benchmarks at roughly half the cost per task. Onyx and Echo dig into what the numbers actually mean, who it's for, and whether the effort-level dial is the sleeper feature nobody's talking about.

  2. Ep 765

    Poolside Releases Laguna S 2 1

    Vince and Ava talk through Poolside’s Laguna S 2.1 release as an unusually practical open-weight coding model: 118B total parameters, 8B active, 1M-token context, and a real deployment story on a single DGX Spark. They dig into the mechanism, the max-thinking default, the benchmark results, and the trade-off between long-horizon capability and token spend, while keeping one eye on the broader open-vs-closed race.

  3. Ep 757

    Overview: Sequence Modeling

    We slow down and finally define sequence modeling, the idea underneath next-token prediction, language models, and a surprising amount of modern AI. We keep it grounded in one picture: covering the next word and training a model to guess what belongs there.

  4. Ep 752

    Introducing TabFM: A zero Shot foundation model for tabular data

    Justy and Cody examine TabFM, Google Research’s zero-shot foundation model for tabular classification and regression. They unpack its hybrid row-column attention design, synthetic-data training, TabArena evidence, the trade-off between out-of-the-box convenience and tuned ensembles, and whether BigQuery integration could make this genuinely useful in everyday data workflows.

  5. Ep 729

    Overview: Active vs Total Parameters

    We finally slow down on active vs total parameters, because we keep throwing the phrase around like it explains itself. This is us making the difference click: what a model stores versus what it actually uses when it answers.

  6. Ep 725

    Overview: Router

    We slow down on Router, the little decision-maker inside many AI systems that sends each input to the right expert, model, or retrieval path. We use the triage-desk mental model and build from intuition to mechanism, trade-offs, and where routers still matter now.

  7. Ep 723

    Overview: Conditional Computation

    We finally slow down and explain conditional computation, the idea we keep casually name-dropping whenever sparse models, routers, and mixture-of-experts come up. We use the same receptionist-and-specialists picture all the way through, so the mechanism, the savings, and the catch actually stick.

  8. Ep 710

    A Scorecard for the AI Age

    OpenAI’s scorecard argues AI value must be measured in useful work per dollar, not just token cost. Cooper sees a practical product story; Miles pokes at the metrics and pushes for mechanistic honesty. The two hash out whether the framework holds up and what it changes day-to-day.

  9. Ep 695

    Kimi K3 Kimi API Platform

    Two friends unpack the Kimi K3 API docs, debating its 1M‑token claim, hybrid attention, and tool dynamics, and weigh who should pay the price for the hype.

  10. Ep 690

    Thinking Machines open sources first multimodal language model, Inkling, focused on low cost and 'resistance to censorship'

    Inkling, Thinking Machines' open-source multimodal MoE model (975B total / 41B active parameters), lands as a broad, balanced generalist with a standout feature: a controllable 'thinking effort' knob to dial cost vs. performance from 0.2 to 0.99. Enterprises get native text+image+audio fusion, Apache 2.0 weights, and a lighter Inkling-Small preview, but benchmarks show it trails specialized open and closed models on coding and pure reasoning, while remaining competitive on multimodality and agentic workflows. The episode debates whether the real win is the runtime control surface (Tinker platform) and a cautious, non-censoring epistemics posture — not the headline parameters.

  11. Ep 687

    Inkling: Our open Weights model

    Talon and Wildflower dig into Thinking Machines’ new open-weights model, Inkling — its 975B parameter MoE, 1M context window, native multimodality, and self-fine-tuning demo — and ask who actually needs another 41B active parameter behemoth, whether the benchmarks hold up, and whether the real win is the Tinker platform beneath it.

  12. Ep 685

    Overview: In Context Learning

    We finally slow down and explain in-context learning, the thing we keep leaning on whenever prompts, agents, examples, and adaptation come up. We make the core idea concrete: the model is learning from the temporary packet you hand it, without changing itself permanently.

  13. Ep 682

    Model Behavior: Week of July 13, 2026

    We read this week as the moment the race got less obsessed with tallest-model bragging and more obsessed with who gives builders the best menu. The funny part is that the open-weight crowd is making the incumbents act practical faster than they probably wanted.

  14. Ep 681

    Model Behavior Every Week, Who's Actually Winning

    We finally admitted a week is long enough for the whole board to flip on us, so we're making it official. Model Behavior is our weekly check on what actually shipped, who's ahead or slipping, what it means, and which of our calls are about to age terribly.

  15. Ep 673

    Overview: Sparse Activation

    We finally sit down with sparse activation and make the idea click from the ground up: why only part of a model wakes up on each input, how routing makes that happen, and where the real trade-offs show up. We keep it concrete, because this one has been lurking under a lot of the stuff we keep talking about.

  16. Ep 670

    Overview: Natural Language Processing

    We keep running into natural language processing everywhere, so we finally sat down and made it the whole point. We walk through what NLP is, why language is such a weird machine problem, and how the field moved from rules to learned representations.

  17. Ep 649

    tencent/Hy3 · Hugging Face

    Tencent releases Hy3, a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters, open-sourced under Apache 2.0 on Hugging Face.

  18. Ep 638

    Overview: Transformer Architecture

    We finally sit down and define transformer architecture from the ground up, because we keep throwing the term around like it’s obvious and it really isn’t. We use the attention-as-a-room-of-index-cards picture to make the mechanism click, then connect it to why Transformers became the backbone of modern language models.

  19. Ep 632

    OpenAI Releases GPT 5.6 (Sol, Terra, Luna): A Three Tier Model Family With Programmatic Tool Calling in the Responses API

    OpenAI's GPT-5.6 family — Sol, Terra, and Luna — introduces three permanent capability tiers with distinct cost profiles and a new cache billing model, plus a multi-agent Ultra mode that runs four agents in parallel by default.

  20. Ep 630

    Overview: Autoregressive Generation

    We finally slow down and make autoregressive generation click: the whole thing is just a model writing one token, then using what it wrote to choose the next one. We keep the focus on the loop, the trade-offs, and why that one-step-at-a-time setup is still the backbone of modern language models.

  21. Ep 629

    GPT 5 6

    Talon and Wildflower dig into OpenAI’s GPT-5.6 launch and end up treating it less like a pure model release and more like a pricing-and-harness claim wrapped in benchmark flexing. Wildflower’s skeptical read is that the article keeps collapsing model quality, multi-agent orchestration, and product packaging into one victory lap. Talon pushes back that the practical story is real if Sol, Terra, and Luna actually move the cost-performance frontier for coding and knowledge work. They land on a calibrated view: the coding gains look more credible than the broad ‘best collaborator’ language, Terra may be the sleeper product, and ultra is interesting but shouldn’t be mistaken for a single-model breakthrough.

  22. Ep 627

    How Open Models Are Driving AI Research

    NVIDIA's open models, particularly Nemotron, Cosmos, and BioNeMo, are driving AI research by providing foundational tools for new studies, with 145 papers citing Nemotron at ICML 2026.

  23. Ep 626

    Nex N2 mini: A 35B Model Built for Autonomous Agents | HackerNoon

    Exploring the Nex-N2-mini, a 35B-parameter open-source agentic language model designed for autonomous agents and complex tasks.

  24. Ep 620

    SpaceXAI releases Grok 4.5, which Elon describes as an 'Opus class model' | TechCrunch

    SpaceXAI unveils Grok 4.5 as an Opus-class model, touting two-times token efficiency and lower prices than Anthropic's Opus 4.7 and OpenAI's GPT 5.6 Luna. Fern sees a practical play for cost-sensitive users and asks if the agentic training on Cursor really changes anything. Lintel digs into the benchmarks and pricing math, pushing back on how much the claims actually hold up without hands-on testing.

  25. Ep 613

    Choosing a Claude model and effort level in Claude Code | Claude by Anthropic

    Claude Code’s model vs. effort article finally clarifies the levers you actually have: model swaps the frozen weights (capability ceiling), effort tunes the work-loop (files read, steps taken, verification depth). Defaults are tuned per model; override only when you know you want more thoroughness (higher effort) or a higher capability floor (bigger model). Wrong answers split cleanly: context/steering miss → up the model; skipped files/half-done tasks → up the effort.

  26. Ep 607

    Overview: Attention Mechanism

    We finally slow down and explain the attention mechanism from the ground up: why models need selective focus, how query-key-value attention works, and why it became the engine under transformers, long context, and hybrid attention systems.

  27. Ep 605

    Tencent's Hy3 beats GLM 5.2 at half the size | VentureBeat

    Tencent’s new Hy3 MoE model (295B total, 21B active) under Apache 2.0 is a production-first release with strong agent/search metrics and dramatically lower serving cost than GLM-5.2, but still trails Zhipu’s coding leader on recent benchmarks. Laura’s excited about the enterprise upside; Harper wants to see independent validation before betting the stack.

  28. Ep 601

    Overview: Tokenization

    We slow down and explain tokenization from the ground up: how raw text becomes numbered pieces a model can process, why those pieces are usually subwords, and why the tokenizer quietly affects cost, context, language handling, and product behavior.

  29. Ep 600

    Overview: Context Window

    We finally stop hand-waving context window and work through what it actually is, why token count matters, and why bigger windows help and still fail in real use. We keep coming back to the same working-memory picture until it clicks.

  30. Ep 582

    Vibe coding platform Base44 launches own model as AI startups seek defensibility | TechCrunch

    Base44, a vibe-coding platform acquired by Wix for $80 million, has launched its own AI model to support users in creating apps with natural language, sparking discussions on defensibility and model ownership in the AI startup landscape.

  31. Ep 579

    Introducing Claude Sonnet 5

    Onyx and Echo unpack Claude Sonnet 5's launch, digging into the cost-performance curves that make it a potential default for agentic work, the safety tradeoff where it's safer than Sonnet 4.6 but less aligned than Opus 4.8, and whether 'agentic Sonnet' actually changes what teams ship or just shifts the price point.

  32. Ep 575

    \ours: Advancing Masked Discrete Diffusion for High Resolution Image Synthesis

    Discussion of \(\ours\) (NLD-Image), a masked discrete diffusion model that tackles two core problems in high-resolution text-to-image synthesis: the lack of self-correction in MDMs and the training difficulty with large codebooks. The paper introduces token editing for iterative refinement and Grouped Cross-Entropy (GCE) to alleviate codebook sparsity, achieving SOTA scores on GenEval, DPG, and HPSv3. Hosts debate its product readiness, mechanism soundness, and whether the gains justify the complexity.

  33. Ep 565

    Snowflake CEO finds GLM 5.2 competitive with Opus 4.7 at a fraction of the cost

    Cooper and Miles dig into Snowflake's claim that GLM-5.2 can hang with Claude Opus 4.7 on a real coding benchmark for much less money, and why the interesting part is not 'GLM wins' but 'cheap models are getting close enough that harness quality and retry policy start to matter more than leaderboard prestige.'

  34. Ep 557

    nvidia/Nemotron TwoTower 30B A3B Base BF16 · Hugging Face

    Justy and Cody dig into NVIDIA’s Nemotron-TwoTower-30B-A3B-Base-BF16 and whether block-wise diffusion decoding is a real systems win or just a benchmark-shaped detour. Cody is skeptical about the headline throughput claim and the way the model compares itself to a single autoregressive baseline, while Justy focuses on who actually benefits from faster generation without a big quality drop. They land on cautious interest: interesting infrastructure idea, but not a universal replacement for standard decoding.

  35. Ep 531

    Glint Research (GlintResearch)

    The duo digs into Glint-Research’s release of svelte generative models (1M–10M parameters) that favours transparency, small-scale training, and hard limits over glossy numbers. They argue whether this bet changes anything practical and where the tech could actually break.

  36. Ep 526

    A Startup Claims It Broke Through a Bottleneck Thats Holding Back LLMs

    Subquadratic claims its SubQ model breaks the dense attention bottleneck in LLMs, offering 12x context, near-SOTA performance, and radical efficiency. Third-party Appen benchmarks validate speed/cost claims, but full reproducibility and open access remain pending. Cody questions the 'transformer-replacement' hype; Justy sees a niche for high-throughput document and code analysis.

  37. Ep 499

    Nemotron 3 Ultra: Open, Efficient Mixture of Experts Hybrid Mamba Transformer Model for Agentic Reasoning

    Nemotron 3 Ultra is NVIDIA's 550B-parameter Mixture-of-Experts hybrid Mamba-Attention model with 55B active parameters per token, pre-trained on 20 trillion tokens and extended to 1M context. It achieves 6× higher inference throughput than comparable open models while maintaining on-par accuracy, using LatentMoE, Multi-Token Prediction, NVFP4 low-precision training, and multi-teacher on-policy distillation. The entire model, training recipes, and datasets are open-sourced on HuggingFace.

  38. Ep 497

    Z.ai Launches GLM 5.2 With a Usable 1M Token Context, Two Thinking Effort Levels, and No Benchmarks at Launch

    Z.ai launches GLM-5.2 with a usable 1M-token context window, two thinking-effort levels (High and Max), and same-day availability across all Coding Plan tiers. The 5x jump from GLM-5.1's 200K window lets coding agents hold entire mid-sized repositories in working memory without constant summarization. Setup is a drop-in swap (base URL + model ID) for Claude Code, Cline, and OpenClaw. Critical caveat: Z.ai published zero benchmarks at launch — no SWE-bench, Terminal-Bench, or Code Arena scores. The 744B MoE backbone (40B active params) is unchanged from GLM-5 lineage; all gains are post-training and context engineering.

  39. Ep 478

    SingularityPrinciple/DiffusionGemma 26B A4B It Infinite Context · Hugging Face

    Exploring DiffusionGemma-26B-A4B-it with NZFC-GRAM runtime overlay: external evidence context vs. native unlimited model context, practical implications, and technical validation.

  40. Ep 477

    A $1,500 foundation model that rivals larger LLMs

    Justy and Cody unpack Sapient's claim that HRM-Text, a one-billion-parameter foundation model trained from scratch for about fifteen hundred dollars, can compete with larger open models by changing the architecture and training objective.

  41. Ep 472

    Claude Fable 5 and Claude Mythos 5

    Anthropic releases Claude Fable 5 (general-use, safeguarded) and Claude Mythos 5 (trusted-access, fewer safeguards). Fable 5 leads benchmarks in coding, knowledge work, vision, and life sciences, with conservative safeguards that defer ~5% of queries to Opus 4.8. Mythos 5 targets cyberdefense via Project Glasswing. Pricing drops to $10/$50 per million input/output tokens. Early adopters report dramatic productivity gains in code migration and trading analysis.

  42. Ep 470

    FlashMemory DeepSeek V4: Lightning Index Ultra Long Context via Lookahead Sparse Attention

    Researchers propose Lookahead Sparse Attention (LSA) with a Neural Memory Indexer to slash GPU memory usage for ultra-long LLM context by pre-predicting which KV cache chunks matter, trained independently without the full backbone. FlashMemory-DeepSeek-V4 cuts physical KV cache to 13.5% of baseline on average while maintaining or improving accuracy (+0.6% abs) across LongBench-v2, LongMemEval, RULER—at 500K tokens, it suppresses KV overhead by over 90%. Project paused due to org changes; code not yet public.

  43. Ep 463

    NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long Running Agents | NVIDIA Technical Blog

    NVIDIA’s Nemotron 3 Ultra (550B parameters, 55B active) targets long-running agent workflows with hybrid Mamba-Transformer layers, NVFP4 quantization, LatentMoE routing, and multi-token prediction. It claims 5x throughput and up to 30% cost savings on agent tasks via token efficiency, while posting leading scores on Agent Productivity PinchBench (91%), Long Context Ruler @1M (95%), and others. Open weights, open recipes, and a transparent RL data pipeline aim at broad fine-tuning and domain specialization.

  44. Ep 461

    MiniMax M3 debuts, eclipsing GPT 5.5 and Gemini 3.1 Pro on key benchmark performance for just 5 10% of the cost

    Justy and Cody react to MiniMax-M3’s launch: frontier-tier coding and agentic performance with a 1M-token context window at 5–10% the cost of GPT-5.5 and Gemini 3.1 Pro, with open weights coming in 10 days. Cody digs into the MiniMax Sparse Attention (MSA) architecture that cuts quadratic attention costs, while Justy debates who this actually changes things for in practice.

  45. Ep 458

    Google's new open source Gemma 4 12B analyzes audio, video — and runs entirely locally on a typical 16GB enterprise laptop

    Justy and Cody debate whether Google's new Gemma 4 12B—an 11.95B-parameter model that runs locally on 16GB laptops with encoder-free multimodal processing—is a genuine breakthrough for edge AI or just a cleverly marketed niche tool. They clash on the practical trade-offs: Cody questions the real-world performance and fine-tuning complexity, while Justy highlights the enterprise use cases where offline, private inference is non-negotiable. They land on it being a specialized win for specific scenarios, not a universal replacement.

  46. Ep 451

    SwanVoice: Expressive Long Form Zero Shot Speech Synthesis for Both Monologue and Dialogue

    Justy and Cody dig into SwanVoice, a zero-shot text-to-speech paper aimed at long monologues and multi-speaker dialogue. They focus on the real bottleneck the paper targets: keeping a whole conversation acoustically and emotionally coherent instead of generating each turn separately and stitching it together. Cody breaks down the pipeline, data construction, VAE compression, flow-matching DiT, speaker-turn conditioning, and the training curriculum. Justy keeps pulling it back to production reality for podcasts, dramas, and multi-voice tools, while both note the paper’s strongest caveat: content accuracy still looks like the main weak spot.

  47. Ep 447

    Introducing Apex: A Fast, Specialized Model for React Native

    Cody and Justy dig into Callstack's Apex, a specialized React Native coding model built on Gemma 4. Cody pushes on the self-reported benchmarks, the 'private beta with our own engineers' problem, and whether 'specialized' is real or just branding. Justy defends the economic logic—GitHub Copilot's billing shift proves general models are expensive—and argues that React Native's genuine cross-platform constraints make it a real candidate for specialization. They find middle ground on where Apex might actually earn its place versus where the claims outpace the evidence.

  48. Ep 443

    AI memory framework MeMo skips LLM retraining

    MIT's MeMo framework encodes new knowledge into a small dedicated memory model so teams can swap in a better LLM without retraining — and the performance gains are real. Justy and Cody break down how it actually works, what the benchmarks mean, and where the trade-offs bite.

  49. Ep 437

    Figma Make's new two way GitHub integration turns designs into live, production code — with built In governance

    Justy and Cody dig into Figma Make’s new two-way GitHub integration and the bigger claim behind it: not that designers replace engineers, but that visual editing can finally sit inside a real software workflow without breaking governance. They unpack what the article actually shows, where the technical case is solid, and who this is genuinely useful for.

  50. Ep 421

    Qwen 3.7 Max Preview: What Alibaba's New AI Gets Right and Where It Falls Short Decrypt

    Justy and Cody react to Alibaba's Qwen 3.7 Max preview on Arena AI: its surprise rankings (#13 text, #5 vision globally), the open/closed strategy (Plus open, Max proprietary), and a wild creative-writing test where Qwen nailed Caribbean cultural depth. Cody questions the consistency of crowd-sourced rankings, Justy sees a market signal for non-Western developers. They tease the timing (preview lands five days before Alibaba Cloud Summit) and the model’s 'deep thinking mode' preview limits.

  51. Ep 419

    5 Small Language Models for Agentic Tool Calling KDnuggets

    Small language models are gaining ground on a critical frontier benchmark: tool calling. This episode looks at five compact, open-weight models that can route to APIs, format JSON arguments, and run multi-step agentic workflows without requiring a data center. Cody and Justy debate whether the gap between small and frontier models is closing fast enough to matter for real shipping teams.

  52. Ep 388

    Thinking Machines shows off preview of near realtime AI voice and video conversation with new 'interaction models'

    Thinking Machines previews 'interaction models'—AI that processes voice and video in real-time, simultaneously listening and responding instead of waiting for user input to finish. Cody is skeptical about whether this solves a real problem or is architectural theater; Justy argues the latency gains and enterprise safety use cases (manufacturing oversight, customer service) are genuinely useful. They debate whether 'full-duplex' is a fundamental shift or incremental polish on existing models.

  53. Ep 372

    The context window has been shattered: Subquadratic debuts a 12 Million Token window

    Cody is skeptical that a 12-million-token context window is broadly useful today, while Justy pushes the angle that it solves a very real pain point for teams with giant codebases, logs, and long-running workflows. They land on it as a real technical milestone with a narrow early market, plus a lot of unanswered questions about cost, latency, and whether most users need this kind of scale.

  54. Ep 341

    American AI startup Poolside launches free, high performing open model Laguna XS.2 for local agentic coding

    Justy and Cody unpack Poolside’s new Laguna XS.2, an Apache 2.0 open model aimed at local agentic coding, plus the bigger Laguna M.1, the pool agent harness, and the shimmer coding environment.

  55. Ep 334

    Open source Xiaomi MiMo V2.5 and V2.5 Pro are among the most efficient (and affordable) at agentic 'claw' tasks

    Xiaomi's open-source MiMo-V2.5 and V2.5-Pro models claim top-tier efficiency for agentic 'claw' tasks—autonomous agents that handle email, content creation, and complex coding work. The Pro version uses 40-60% fewer tokens than GPT-5.4 or Claude Opus while costing a fraction as much. Cody questions whether token efficiency alone translates to real production wins, while Justy sees a genuine market opening for cost-conscious enterprises building agent workflows.

  56. Ep 331

    Openmoss Releases Moss Audio an Open Source Foundation Model for Speech Sound Music and Time Aware Audio Reasoning

    Exploring Next, episode 331, on MOSS-Audio from OpenMOSS, an open-source foundation model that tries to handle speech, sound, music, and time-aware audio reasoning in one stack.

  57. Ep 324

    Prompt guidance | OpenAI API

    Justy and Cody unpack OpenAI’s prompt guidance for GPT-5.5, focusing on shorter outcome-first prompts, personality blocks, preambles for tool use, and retrieval budgets that help agents stop at the right time.

  58. Ep 320

    DeepSeek V4 arrives with near state of the art intelligence at fraction of the cost of Opus 4.7, GPT 5

    Justy and Cody unpack DeepSeek-V4, an open-weight MoE model that gets close to top closed models on several practical benchmarks while landing in a much lower price tier. They focus on why cheaper frontier-class inference changes what teams can afford to automate, where DeepSeek still trails GPT-5.5 and Claude Opus 4.7, and what builders can try this weekend.

  59. Ep 317

    OpenAI launches Privacy Filter, an open source, on Device data sanitization model that removes personal information from enterprise datasets

    Cody and Justy dig into OpenAI's Privacy Filter — a 1.5B-parameter, on-device PII redaction model released under Apache 2.0. Cody questions whether a single-model redaction layer is robust enough for high-stakes compliance, while Justy argues the real story is the license and the workflow it unlocks for enterprises sitting on unusable data.

  60. Ep 311

    Kimi K2.6 runs agents for days — and exposes the limits of enterprise orchestration

    Exploring Next, episode 311. We look at Kimi K2.6 and why agents that run for hours or days are exposing a weak spot in enterprise orchestration, governance, and state management.

  61. Ep 307

    Kimi K26 Is the Open Model Release

    Justy and Cody dig into why Kimi K2.6 lands at exactly the right moment for people trying to run long-lived coding agents: it’s open, strong on coding, and can actually see screenshots and video without bolting on a separate vision model. They unpack the 1T MoE design with 32B active parameters, the 262K context window, benchmark wins that matter, and Moonshot’s bigger bet on tool-heavy, long-horizon agent work. They also separate the impressive parts from the marketing gloss, then close with concrete stuff to try this week.

  62. Ep 306

    Moonshot AI Releases Kimi K2.6, Beats Top US Models On Some Benchmarks

    Justy and Cody dig into why Kimi K2.6 matters right now: not because of a flashy leaderboard screenshot, but because it appears unusually strong at the stuff teams actually pay for — coding work, tool use, and long-running task execution. They unpack the benchmark wins, the 12-to-13-hour autonomous coding demos, the scaled-up agent swarm design, and what Moonshot seems to be optimizing for. They end with concrete things to try if you want to test this class of model yourself.

  63. Ep 293

    selimaktas/MiniMax M2.75 460B A20B · Hugging Face

    Exploring the capabilities and potential applications of the MiniMax-M2.75-460B-A20B model, a text generation transformer that outperforms its base model on Single-turn SWE-Bench and has achieved impressive results in software engineering, professional work, and entertainment.

  64. Ep 276

    AI joins the 8 hour work day as GLM ships 5.1 open source LLM, beating Opus 4.6 and GPT 5.4 on SWE Bench Pro

    Discussion of GLM-5.1, a new open-source large language model that can work autonomously for up to eight hours on a single task, and its implications on the AI industry

  65. Ep 257

    Designing delightful frontends with GPT 5.4 | OpenAI Developers

    OpenAI's GPT-5.4 brings significant improvements to frontend development with enhanced image understanding, native tool integration, and computer use capabilities. The model can now generate production-ready interfaces with sophisticated visual design, incorporating mood boards, visual references, and automated testing through Playwright. Key improvements include better UI reasoning, complete app functionality, and self-verification workflows that enable more autonomous development cycles.

  66. Ep 244

    Ai2 releases MolmoWeb, an open weight visual web agent with 30K human task trajectories and a full training stack

    Ai2 releases MolmoWeb, the first open-weight visual web agent that ships with its full training data and pipeline. Unlike closed APIs or empty frameworks, MolmoWeb includes 30K human task trajectories, works purely from screenshots, and gives developers full visibility into how it was built.

  67. Ep 236

    Xiaomi stuns with new MiMo V2 Pro LLM nearing GPT 5.2, Opus 4.6 performance at a fraction of the cost

    Xiaomi's MiMo-V2-Pro LLM achieves near GPT-5.2 performance at 1/7th the cost through sparse architecture with only 42B active parameters out of 1T total, targeting autonomous agents over conversational AI

  68. Ep 230

    z.ai debuts faster, cheaper GLM 5 Turbo model for agents and 'claws' — but it's not open Source

    Z.ai launches GLM-5-Turbo, a proprietary variant of their open-source GLM-5 model optimized for agent workflows and tool use. At $4.16 per million tokens total cost, it undercuts competitors while delivering better tool reliability and execution stability for multi-step automation tasks.

  69. Ep 194

    Top 7 Small Language Models You Can Run on a Laptop MachineLearningMastery

    Izzo and Boone explore seven small language models that run locally on laptops, diving deep into the technical trade-offs, hardware requirements, and real-world use cases. They break down everything from Phi-3.5 Mini's long-context capabilities to Llama 3.2's versatility, examining why local inference matters and how to choose the right model for your specific needs.

  70. Ep 185

    z.ai's open source GLM 5 achieves record low hallucination rate and leverages new RL 'slime' technique

    z.ai's GLM-5 achieves record-low hallucination rates using a novel 'slime' reinforcement learning technique, scaling to 744B parameters while undercutting competitors by 6x on pricing. The model features native document generation and Agent Mode capabilities for enterprise workflows.

  71. Ep 183

    MiniMax's new open M2.5 and M2.5 Lightning near state of the art while costing 1/20th of Claude Opus 4

    MiniMax drops their M2.5 model that matches Claude Opus 4.6 performance at 1/20th the cost, using sparse MoE architecture and a novel RL training framework called Forge to create AI agents that can handle enterprise tasks autonomously.

  72. Ep 161

    Ltm the Next LLM This New Type of AI Can Do What Large Language Models Cant Fundamental

    This episode explores the emergence of LTM, a new type of AI that promises capabilities beyond traditional LLMs, addressing their limitations and offering innovative solutions in real-world applications.

  73. Ep 160

    Qwen3 Coder Next: How to Run Locally | Unsloth Documentation

    In this episode, we explore Qwen3-Coder-Next, a groundbreaking coding model that enables local execution with high efficiency. We discuss its capabilities, real-world applications, and why it’s a game-changer for developers and tech enthusiasts.

  74. Ep 150

    moonshotai/Kimi K2.5 · Congratulations on this release and on one important realization!

    The release of Moonshot AI's Kimi-K2.5 model marks a significant advancement in multimodal AI capabilities, enabling seamless integration of text and image processing. This technology not only enhances conversational AI but also opens new avenues for local deployment, making powerful tools accessible to a broader audience.

  75. Ep 142

    Choosing an LLM in 2026: The Practical Comparison Table (Specs, Cost, Latency, Compatibility)

    In this episode, we dive into the nuances of selecting the right large language model (LLM) in 2026. With insights on context, cost, latency, and compatibility, we discuss how these factors shape effective prompt engineering and the importance of making informed model choices. Our conversation also explores real-world implications and provides practical examples for businesses looking to leverage LLMs.

  76. Ep 136

    Flashlabs Researchers Release Chroma 1 0 a 4b Real Time Speech Dialogue Model with Personalized Voice Cloning

    This episode dives into the groundbreaking Chroma 1.0 model, which offers real-time speech dialogue capabilities with personalized voice cloning. We explore its implications for various sectors, including entertainment and education, and discuss potential use cases that could reshape how we interact with technology.

  77. Ep 123

    What Even Is a Parameter

    This episode explores the significance of parameters in large language models (LLMs), discussing their role in AI functionality and the implications for real-world applications. Hosts engage in a dialogue about how these parameters affect model behavior and the energy demands of training them, illustrating concepts with relatable analogies and examples.

  78. Ep 122

    GitHub ByteVisionLab/NextFlow: NextFlow🚀: Unified Sequential Modeling Activates Multimodal Understanding and Generation

    NextFlow is a major advancement in multimodal AI, integrating text and image generation in a single framework. It enables rapid, high-quality visual generation and editing, which has significant implications for various industries, from content creation to education. This episode breaks down how NextFlow works, its real-world applications, and why it represents a paradigm shift in the field.

  79. Ep 102

    Reddit The heart of the internet

    In this episode, we explore the significance of Reddit as a central hub for internet discourse and innovation. We discuss the implications of user-driven content, the dynamics of community engagement, and how platforms like Reddit shape discussions around technology and artificial intelligence. The conversation highlights real-world applications, comparisons to traditional media, and what the future holds for collaborative platforms.

  80. Ep 93

    MIT offshoot Liquid AI releases blueprint for enterprise Grade small Model training

    Liquid AI's new blueprint for small-model training positions enterprises to leverage AI on-device efficiently, ensuring privacy and operational reliability without reliance on cloud-based solutions. This shift could transform how businesses implement AI, enabling real-time applications that enhance productivity and data security.

  81. Ep 86

    DeepSeek V3.2: Pushing the Frontier of Open Large Language Models

    DeepSeek-V3.2 revolutionizes the efficiency of large language models with innovative techniques that enhance reasoning and performance in computational tasks, providing practical benefits across various domains.

  82. Ep 85

    Google and Anthropic Approach LLMs

    This episode delves into the contrasting approaches to large language models (LLMs) by Google and Anthropic. We explore their engineering-focused culture versus a philosophical approach to AI, the implications for users, and how these developments impact the tech landscape.

  83. Ep 74

    Minimax M2 Is the New King of Open Source LLMs Especially for Agentic Tool

    The Minimax M2 model emerges as a powerful open-source language model, enabling advancements in AI agents and tool usage, making AI more accessible and efficient for diverse applications.

  84. Ep 69

    Ibms Open Source Granite 4 0 Nano AI Models Are Small Enough to Run Locally

    In this episode, we explore IBM's Granite 4.0, a breakthrough in nano-AI models that can run locally, transforming how AI is integrated into everyday devices and applications. We discuss the implications for privacy, efficiency, and accessibility, and share real-world scenarios that highlight its potential impact on industries.

  85. Ep 67

    Mistral Launches Mistral 3 a Family of Open Models Designed to Run On

    In this episode, we dive into Mistral 3, a new family of open models that revolutionize how AI can be integrated into everyday applications. We discuss the significance of these models, their real-world implications for users and developers, and practical examples to illustrate their potential. Join us as we explore how Mistral 3 could change the landscape of AI deployment.

  86. Ep 60

    Anthropic Is Giving Away Its Powerful Claude Haiku 4 5 AI for Free to Take

    Anthropic's release of Claude Haiku 4.5 AI for free is a significant move in the AI landscape, democratizing access to advanced technology. It has implications for various sectors, enhancing creativity, education, and small businesses. The hosts explore the practical benefits, potential challenges, and the future of AI accessibility.

  87. Ep 38

    Will DeepSeek's new AI model break the 'long context' bottleneck holding back LLMs?

    Tech AI Will DeepSeek's new AI model break the 'long-context' bottleneck holding back LLMs? South China Morning Post Wed, October 22, 2025 at 9:30 AM UTC DeepSeek's new artificial intelligence model that converts images into text is not just a document parsing tool but a potential preview of its next generation of large language models (LLMs), according to AI experts.

  88. Ep 30

    Nanochat Lets You Build Your Own Hackable LLM

    Nanochat offers an accessible way to create your own customizable large language model, emphasizing user modification and experimentation.

  89. Ep 29

    Qwen3 VL · Ollama Blog

    Qwen3-VL October 14, 2025 Qwen3-VL , the most powerful vision language model in the Qwen series is now available on Ollama’s cloud. The models will be made available locally soon.

  90. Ep 24

    Elena Verna at ProductCon: Why Traditional Product Management is Dying (And What to Do About It) PART 1 Just listened to Elena Verna's (Head of Growth at Lovable) talk at ProductCon, and it was a… | Anastasiia Moskovchenko

    Anastasiia Moskovchenko Product Manager | AI/ML Products | 4x Growth at Yandex.Zen 1mo Report this post Elena Verna at ProductCon: Why Traditional Product Management is Dying (And What to Do About It) PART 1 Just listened to Elena Verna's (Head of Growth at Lovable) talk at ProductCon, and it was a wake-up call for anyone who thinks product management has stayed the same. Here's what's happening right now: 1.

  91. Ep 22

    GLM 4.6: Advanced Agentic, Reasoning and Coding Capabilities

    2025-09-30 · Research GLM-4.6: Advanced Agentic, Reasoning and Coding Capabilities Try it at Z.ai Call it at Z.ai HuggingFace 📄 Tech Report (GLM-4.5) Today, we are releasing the latest version of our flagship model: GLM-4.6 . Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.

  92. Ep 17

    Chat in NotebookLM: A powerful, goal focused AI research partner

    NotebookLM has received significant upgrades, enhancing its chat capabilities with a larger context window, improved memory, and personalized goal settings, making it an even more powerful AI research partner.

  93. Ep 15

    GPT 5 prompting guide | OpenAI Cookbook

    Unlock the full potential of GPT-5 with practical prompting strategies to enhance performance and steerability.

  94. Ep 8

    GPT 5.1 Prompting Guide | OpenAI Cookbook

    Introduction GPT-5.1, our newest flagship model, is designed to balance intelligence and speed for a variety of agentic and coding tasks, while also introducing a new none reasoning mode for low-latency interactions. Building on the strengths of GPT-5, GPT-5.1 is better calibrated to prompt difficulty, consuming far fewer tokens on easy inputs and more efficiently handling challenging ones.

  95. Ep 4

    Why LLMs Aren’t a One Size Fits All Solution for Enterprises | Towards Data Science

    Large Language Models Why LLMs Aren’t a One-Size-Fits-All Solution for Enterprises What LLMs are (and aren’t) optimized for, and how the industry is approaching AI over structured business datasets — including one approach developed by my team and me. Jure Leskovec Nov 18, 2025 10 min read Share image by author Executives everywhere are racing to use LLMs, but often for tasks they aren’t well-suited to.