Hacker Newsnew | past | comments | ask | show | jobs | submit | mikenew's commentslogin

Apple is on the side of privacy if it serves their marketing. Which is why they would rather build and normalize CLIENT SIDE CONTENT SCANNING so they can continue to market iCloud as "secure and private".

If Apple's interests sometimes align with ours then great. I'll take it. But don't attribute to this ~5 trillion dollar company some kind of altruism.


> If Apple's interests sometimes align with ours then great.

Altruism or not, I've found that their interests almost always align with mine when it comes to privacy.


I've found that Apple's interests are pretty fickle, concerning real-world security: https://www.securityweek.com/apple-suddenly-drops-nso-group-...

> Apple has abruptly withdrawn its lawsuit against NSO Group, citing increased risk that the legal battle might unintentionally reveal sensitive vulnerability data and difficulties in acquiring essential information from the spyware vendor.

> In a court filing Friday, Apple said continuing the lawsuit now poses “too significant a risk” of exposing the anti-exploitation and threat intelligence efforts needed to fend off the very adversaries involved in the legal dispute.

> “When it filed this lawsuit nearly three years ago, Apple recognized that it would involve sharing information with third parties. However, developments since then have reshaped the risk landscape associated with sharing such information,” the Cupertino device maker said.

I mean, iOS has never been open source so Apple has always practiced security through obscurity. It doesn't seem unreasonable that they'd be concerned about this. Idk, seems like a nothingburger but maybe I'm wrong.


https://knowyourmeme.com/photos/1307138-we-did-it-patrick-we...

Dropping the case is great for Apple's obscurity, but terrible for enforcing the security of iOS users that are still vulnerable to Pegasus malware. Now NSO Group is not threatened or deterred, which is the worst of all worlds for iOS security.


Yeah at 120hz each frame is 8ms. So you're only missing a single frame around 30% of the time.

I certainly want my latency as low as I can get it. But I'm pretty skeptical that anyone is truly feeling the difference of a couple ms.


This[1] study suggests image framerate does not significantly affect our ability to perform predicted moves accurately in time, above a certain minimum which is lower than 24 Hz.

Moving the mouse and then pressing a button, or pressing one key and then another, are both cases of predictive moves.

So it seems one shouldn't rely on framerate when talking about the limits of sensing input lag, at least in general.

[1]: https://jov.arvojournals.org/article.aspx?articleid=2213289


Try musicians that are used to playing extremely high speed semihemidemisemiquavers.

We notice latency. Neil Peart could almost get sample-precise timing, he was so godlike.


Humans are good at predicting stuff, and they can notice when the expected doesn't happen.

But I wouldn't necessarily say that people can notice it everywhere in every state of mind. The medium, context etc all matter a lot.


Current top 5 played games on the Steam Deck are: Slay the Spire 2, The Binding of Isaac, Dave the Diver, Stardew Valley, and Baldur's Gate 3. BG3 is the only one you can consider AAA, but Larian is hardly an EA or Blizzard.

The big AAA studios recycle the same formulas and push graphical quality (mostly downstream of Unreal Engine improvements) because they are risk averse. Similar to big comic-book-hero films. It's not what consumers want and it's increasingly starting to show. AAA is struggling badly while indie has been on an absolute tear the past few years.


On top of that, the seemingly endless remakes... Some are praise worthy others not so much. I got tired of Resident Evil Remakes, they are just recycling the same formula over, and over, and over again... Innovation has left a long time a go. Other example is the recent Assassin's Creed Black Flag "Resync" (which is a remaster with some things on top of it), I was planing on buying it, but Jesus, Ubisoft couldn't let it pass: their launcher is still there, all DLC's together cost more than the base game (and why have DLC in the first place?), the menu is has an ad for Assassin's Creed Shadows which, etc.

I'm currently doing a course (I think it's equivalent to an "Associate degree" in USA) in Game Development. I feel demotivated to take part in this industry, and I think that, at least for me, it should be just on and off gig to do in my spare time, perhaps with other enthusiasts too, but I don't plan to pursue it.


Well, because most recent AAA titles would be miserable to play on a handheld.


> Current top 5 played games on the Steam Deck are...

I think this speaks more to the type of games people want in their pocket. Games you can casually pick up, put down.


Top-5 played games off the steam deck are:

Counter-Strike 2 (Valve)

DOTA 2 (Valve)

Palworld (Pocketpair)

PUBG: battlegrounds (PUBG corporation/Krafton Inc)

TBH: Task Bar Hero (obviously not AAA)

I guess Valve would half-qualify as AAA? The top played true AAA game would be, imho "Grand Theft Auto V Enhanced", at the 18th place. Perhaps Ubisoft with a Assasin's creed title in the top sellers probably should be higher in the playing list.

But I think it's fair to say AAA doesn't hold the steam store anywhere near as tightly as it holds the consoles.

https://store.steampowered.com/charts/mostplayed

https://store.steampowered.com/charts/topselling/US


Arcan is such a fascinating project that I've never quite managed to get my head around.


That was my experience a couple months ago, and until someone shows me real evidence of something valuable they've made with it, I'm not wasting my time on this stuff again.


Authenticity requires vulnerability and that's not something Apple can do.


GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all.

Some people seem to agree and some don't, but I think that indicates we're just down to your specific domain and usage patterns rather than the SOTA models being objectively better like they clearly used to be.


It seems like people can't even agree which SOTA model is best at any given moment anymore, so yeah I think it's just subjective at this point.


Perhaps not even necessarily subjective, just performance is highly task-dependent and even variable within tasks. People get objectively different experiences, and assume one or another is better, but it's basically random.


Unless you're looking at something like a pass@100 benchmark, the benchmarks are confounded heavily by a likelihood of a "golden path" retrieval within their capabilities. This is on top of uncertainties like how well your task within a domain maps to the relevant test sets, as well as factors like context fullness and context complexity (heavy list of relevant complex instructions can weigh on capabilities in different ways than e.g. having a history where there's prior unrelated tasks still in context).

The best tests are your own custom personal-task-relevant standardized tests (which the best models can't saturate, so aiming for less than 70% pass rate in the best case).

All this is to say that most people are not doing the latter and their vibes are heavily confounded to the point of being mostly meaningless.


The pass@100 is such a weird critique angle that is surprisingly mainstream; guess what, no one cares if the correct answer is in the top 100, it needs to be the top 1. A model with a better answer in the top 1 is a better model, full stop.


This. Plus if you want to even attempt measuring real 'intelligence' you want to run a neuro-symbolic, de-lexicalized benchmark (e.g. DL-ReasonSuite, SoLT, GSM-Symbolic) - which none of the providers releasing new models showcase.


>just performance is highly task-dependent and even variable within tasks. People get objectively different experiences, and assume one or another is better, but it's basically random.

You are right that this is not exactly subjectivity, but I think for most people it feels like it. We don't have good benchmarks (imo), we read a lot about other people's experiences, and we have our own. I think certain models are going to be objectively better at certain tasks, it's just our ability to know which currently is impaired.


SOTA models war is the new console war.

But more seriously, I can't help but be amused by how emotionally invested in their AI brand of choice people are getting.


AI is a complete commodity

One model can replace another at any given moment in time.

It's NOT a winner-takes-all industry

and hence none of the lofty valuations make sense.

the AI bubble burst will be epic and make us all poorer. Yay


Staying power is probably the most important factor, which is why I'm thinking Google eventually takes the crown.


They might be converging somewhat. The ultimate limiting factor is training data. Eventually I think they will converge and then the competition will be on memory and compute efficiency, with the best being the smallest maximally capable model.


And the subjectivity is bidirectional.

People judge models on their outputs, but how you like to prompt has a tremendous impact on those outputs and explains why people have wildly different experiences with the same model.


I had one occasion where GLM 5.1 did about 95% of the implementation that I needed but couldn't progress form there. And Codex (free quota) solved the remaining 5% on the spot. I'm super happy with both. I don't touch anything Anthropic with a 10 foot pole.


>GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all.

GLM 5.1 is pretty good but there are some "buts".

They hiked the prices 2 times this year. I subscribed to the pro coding plan just before the last hike. At the start of the year, they had only 5 hours quota and no weekly quota. And I hit the weekly quota hard. I can't upgrade the subscription to get a higher weekly quota because they jacked up the prices a lot recently.

My $30 subscription costs now $72. Previously was $15. Max was $49,then $80 and now $160.


What hardware do you run it on? Trying to consider the cost of subscription + API vs new HW..


I used GLM 5.1 and it was bad, I have no clue why people claim it is good


The value in Claude Code is its harness. I've tried the desktop app and found it was absolutely terrible in comparison. Like, the very nature of it being a separate codebase is already enough to completely throw off its performance compared to the CLI. Nuts.


> The value in Claude Code is its harness

If this was the case then Anthropic would be in a very bad spot.

It's not, which is why people got so mad about being forced to use it rather than better third party harnesses.

Pi is better than CC as a harness in almost every respect.


Anthropic limiting Claude subs to Claude code is what pushed me away in the end because I wanted to keep using Pi.


Just sign up for an AWS account and use the Anthropic models through Bedrock which Pi can use.


API costs are really high compared to subs.


Then you aren't the target market.


Why use tricks to support a company that is hostile to your use case?


What advantage are you saying this has compared to just directly going through the Anthropic provider? They are the same price.


Can you enumerate why?


- Claude Code has repeatedly had enormous token wastage bugs. Its agent interactions are also inefficient. These are the cause of many of the reports of "single prompt blew through 5-hour quota" even though it's a reasonable prompt.

- It still lacks support for industry standards such as AGENTS.md

- Extremely limited customization

- Lots of bugs including often making it impossible to view pre-compaction messages inside Claude Code.

- Obvious one: can't easily switch between Claude and non-Claude models

- Resource usage

More than anything, I haven't found a single thing that Pi does worse. All of it is just straight up better or the same.


I thought the desktop app used the cli app in the background?


I feel like it's Sonnet level for implementation, but not matching up to Opus for planning.

But I agree it's close enough that it's worth using heavily. I've not cancelled my Claude Max subscription, but I've added a z.ai subscription...


My combo is codex and claude basic subscription for planing the hard tasks (if any) opencode with GLM 5.1 (z.ai coding plan) for the actual coding.

opencode is awesome I don't miss cluade or codex cli at all, and the z.ai plan is way more generous in compression.

I was lucky to subscribe to z.ai coding plan pro when it costed 30$/month, I was surprised now it costs 70$/month.

In case anyone wants to subscribe to z.ai with 10% discount [1] * here is the credit campaign rules * [2]

- [1] https://z.ai/subscribe?ic=MW6H74HAZ0

- [2] https://docs.z.ai/devpack/credit-campaign-rules


Hmm

Will try it out. Thanks for sharing!


What is your workflow? Do you use Cursor or another tool for code Gen?


I use Opencode, both directly and through Discord via a little bridge called Kimaki.

https://github.com/remorses/kimaki


I've found GLM 5.1 to be extremely good, as long as you keep the context under 100k or so. But I do agree; the actual service is rough. It's often unusable during peak hours.

I imagine they're trying to slow down the acquisition of new customers because they're so overloaded. But yeah at double the price it doesn't seem worth it anymore.


I can't find an announcement on this, but here's the archived subscription page from a couple days ago: https://web.archive.org/web/20260410092340/https://z.ai/subs...

The Max plan has doubled, and the Lite and Pro plans have more than doubled. No change in usage limits.


Even if this is largely due to a change in how PCs in China are being counted, it's still amazing to watch Linux usage continue to climb like this.

It's really the only opposing force to Microsoft's enshittification of Windows.


Linux’s ecosystem has also improved significantly over the past two years, especially in China. Due to the influence of “Xinchuang” (that is, domestically produced Linux rebranded under another shell), many Chinese desktop applications have been reworked in the past couple of years, switching from Windows-specific tech stacks to cross-platform ones—mostly Electron, basically browser wrappers—and now support the Linux platform. The commonly used software is basically all there.

In addition, the development of LLMs has greatly lowered the barrier to using the Linux command line. Problems that used to take a full day to solve can now be handled easily by anyone who can write a prompt—just ask, copy, and paste. This has even made Windows’ command line unfriendly by comparison, despite its own major improvements in recent years, turning it into a significant drawback.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: