Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Asteria and Draco discuss the launch of Gemini 3.8 Flash and 3.8 Flash Cyber, focusing on the 'work harder' reasoning approach and the specialized cybersecurity capabilities for trusted defenders.
Transcript
Asteria I mean... three Flash releases in six weeks. Google is just... they're just throwing the kitchen sink at the workhorse tier right now.
Draco Right? It's almost comical. Like, do they even have a versioning strategy anymore, or is it just... whatever is ready on Tuesday?
Asteria Exactly. But look, 3.8 Flash just landed, and it's actually kind of a big deal for anyone building agents. It's keeping that same introductory price as 3.7... seventy-five cents per million input tokens.
Draco Mm-hm.
Asteria Which is basically a signal that they want this to be the default for every agentic loop. By the way, how's your week going? You still staring at those latency spikes?
Draco Ugh, yeah. It's been a long Wednesday. I'm basically just babysitting a cluster at this point. But anyway... the interesting part here isn't the price, it's the 'diligence' claim.
Asteria Right, the 'works harder' bit.
Draco Yeah. They're explicitly saying 3.8 Flash executes extra reasoning steps and calls tools iteratively on complex tasks. It's not just a weight update; it's a change in how the model handles effort. It might actually use MORE tokens to get the right answer.
Asteria Oh, interesting.
Draco Exactly. Which is a huge trade-off. If you're optimizing for cost, you might actually prefer 3.7. But for long-horizon stuff... like the Deep S W E benchmarks... 3.8 is apparently beating out models that are way larger and more expensive.
Asteria See, that's the product win for me. If I can solve a complex engineering problem end-to-end for a fraction of the cost of a frontier model, I don't care if it uses a few extra tokens to get there. That's just... that's a massive win for the dev experience.
Draco Sure, if the reliability is actually there. I'm always wary of those 'autonomous solving' claims until I see the failure modes. But... the Cyber variant is where it gets weird.
Asteria Right, the Fairwind Program.
Draco Yeah. 3.8 Flash Cyber. It's restricted to 'trusted defenders.' They're claiming frontier-level performance in vulnerability discovery, and they've got this seventy percent success rate on an internal benchmark across twenty languages.
Asteria Wait, seventy percent?
Draco Yeah. And on C W E Bench, it's basically on the Pareto frontier for patching. The Chrome team says it's producing two point six times more correct patches than the best commercial models. That's... actually a significant leap.
Asteria I love that they're prioritizing patching over exploitation. It's such a pragmatic move. Like, 'here is a tool to fix your house' instead of 'here is a tool to break into one.' It makes the safety argument a lot easier to sell.
Draco It does. Though, I'll admit, the 'trusted defender' gate is a classic move. It's a way to ship a high-capability model without the PR nightmare of an automated hacking tool.
Asteria Oh, come on, Draco. You're just being a pessimist. This is a genuine utility for critical infrastructure. Imagine finding a foundational vulnerability in two hours that usually takes months... that's not just a benchmark, that's a real-world save.
Draco Okay, fair. I'll give you that one. It's a real use case. I just... I don't know, I'm still waiting for the day we stop calling everything a 'frontier' model just because it's the newest one in the folder.
Asteria Stop it. You're just grumpy because of your cluster. But seriously, if you're building in AI Studio or Android Studio, you can just swap to 3.8 Flash now. It's already live.
Draco I'll stick to my latency spikes for now.
Asteria You're hopeless. Anyway, I'm going to go see if I can actually get a 3D castle built with a single prompt. Talk later, Draco.