Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
[dupe] Gemini 3.8 Flash and 3.8 Flash Cyber (blog.google)
76 points by simonsarris 5 hours ago | hide | past | favorite | 19 comments
 help



They've interestingly left out any mention of speed.

I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.

Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.

1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...


It depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.

Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?

I guess it might be relative, but switching from VertexAI endpoint to OpenRouter was like 2-3x faster for us.

Hopefully before they release 4.0 Flash we will finally get Gemini 3.5 Pro.

more likely 4 pro will be released pretty soon instead, since pre-training for 4 began in late July

https://x.com/OfficialLoganK/status/2079594867161022817


> 4.0 Flash we will finally get Gemini 3.5 Pro

Nah, we'll just get the 4.0 Pro Preview.


Everyone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.

Supposedly Fable 5.1 is better, but I haven't tried it yet. I've run into the same thing with mundane work that is barely bio/cyber adjacent.

Re: Chinese models, even if the model itself isn't censored, some of the big model providers have guardrails now that you can't exceed, which somewhat defeats the purpose.


"Uncensored" means "weights modified to remove refusals". Abliterated. Providers do not serve such models, at least not frontier-grade. You have to run the weights yourself. For Kimi K3, this is about $60/hour for hardware rental. But you can have about 100 sessions simultaneously.

And yes, Fable 5.1 has the same refusal rate, and significantly nerfed reasoning.



3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...


Who coined the phrase "cyber" for security related things lol. It's so 1999.

Hasn't the field been called "cybersecurity" since... forever?

Cyber is more of an early 1990's thing, and I have no issue with it unlike most in the tech field. I feel like it dropped off in the late 90's and early 00's but made a comeback as hacking became a mainstream security issue.

I assure you that in 1999 "cyber" meant something very different

Amusingly, "cyber" comes from the word "kubernetes"!


A/S/L?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: