Ep 605 News 2:58 w/ Laura & Harper

Tencent's Hy3 beats GLM 5.2 at half the size | VentureBeat

Tencent’s new Hy3 MoE model (295B total, 21B active) under Apache 2.0 is a production-first release with strong agent/search metrics and dramatically lower serving cost than GLM-5.2, but still trails Zhipu’s coding leader on recent benchmarks. Laura’s excited about the enterprise upside; Harper wants to see independent validation before betting the stack.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/605"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 605 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script Mistral Small 4 119B 2603 Voice Cartesia TTS

Transcript

Laura Tencent just shipped Hy3: a 295-billion-parameter MoE, Apache 2.0, half the size of GLM-5.2—and the headline says it wins everywhere except coding. Okay, that is such an Exploring Next take.

Harper Because the coding numbers are still GLM-5.2 ahead? The article admits that.

Laura Yeah, but Hy3’s beating GLM on search, tool orchestration, and long-context retrieval—by nontrivial margins. And crucially, it’s Apache-licensed, so no EU/UK/Korea restrictions.

Harper Right. So legal hurdles gone, but the benchmarks still need third-party eyes.

Laura Come on. They rebuilt the pipeline in ten weeks with 50 internal teams, and the behavior shifted—hallucination down from twelve-and-a-half percent to five-point-four, commonsense errors halved, multi-turn issues cut in half. That’s not nothing.

Harper Internal metrics. Everything’s internal except the architecture numbers.

Laura Okay, but the architecture is the boring part—295B total, 21B active, top-8 routing over 192 experts, MTP layer, 256K context. The real move is the production reliability report disguised as a model card.

Harper Production reliability is the product pitch. I see that. But GLM-5.2 still owns coding—84-point-two on SWE-bench Verified versus Hy3’s seventy-eight. That’s not rounding error.

Laura Hy3’s three-twenty-four on BrowseComp and ninety-one on DeepSearchQA—those are search-and-agent numbers GLM-5.2 isn’t even competing in. And Harper, Hy3’s footprint is under three hundred gigabytes in FP8. GLM-5.2 is seven-forty-four gigs. You can run Hy3 on a single H20-3e instead of an eight-card H200 cluster.

Harper Deployed against export-compliant silicon—nice touch. Eight H20-3es fit the legal envelope, and it still runs fine on H100s everyone else has.

Laura Exactly! So for teams that care about agent harness, search pipelines, or reliability-sensitive apps—where GLM-5.2’s coding crown doesn’t buy them anything—Hy3 looks like the pragmatic open-weight choice.

Harper Pragmatic if the numbers hold up when someone else runs them. Which, fun fact, hasn’t happened yet.

Laura Ugh, Harper. Independent verification is always ‘coming soon.’ Meanwhile, procurement teams just need a license that won’t get them audited into oblivion.

Harper Sure. So we’ll call this the ‘don’t migrate your stack yet’ model—excellent license, promising infra, but wait for the independent eval before betting the farm.

Laura Fine. I’ll take the win that Apache-2.0 Hy3 exists—and it’s free on OpenRouter for two weeks. That’s enough to kick the tires.

Harper Two weeks is plenty of time to spin up a vLLM cluster and see if it hallucinates less than the marketing says.