Hacker Newsnew | past | comments | ask | show | jobs | submit | suriyaG's commentslogin

I was quite inspired by jev-like and built jev compatible layer built on top of gemma.

on my local benchmarks and my personal problems gemma works very well. but training a new attention head for better results in narrow domains should be much easier.


this is a fantastic metaphor. I'm going to borrow it.

> exchange value in the captitalist political economy.

how can one exist in modern society without being an exchange value in the capitalist political economy?

linux is where it is because it is able to be a critical part of said capitalist political economy. even if you want to be a FOSS contributor all your life, you need someone to figure out the "economy" part of it.

food is a false equivalence in this case.


I've taken a few law classes and legal law is frustratingly hard to interpret. I shudder to think what the LLM would end up doing to "follow all relevant laws"

look these up for a fascinating weekend read:

- Beavers and Capybaras are Fish

- Bees are Fish

- Carrots are fruits

- Tomatoes are Vegetables

- X-men are not human


this was me in 2015 in India. only because the inspector wouldn't see anything other than a 500 rupees note.

not sure why you're getting downvoted. but if you're building even very popular applications, it is quite easy to see how LLMs are very well suited for this type of consistency job.

- they follow instructions quite well.

- are tireless at doing mechanical ports between languages and frameworks.

- Can understand a new ecosystem quite well.


Then why aren't people doing that? Microsoft has as much access to cheap tokens as any software company around, right, so you'd think they would be leading the way.

I'm sure these things are in the plans from upper management.

we'll know when the next layoffs hit.


they probably are but have some patience.

Keep in mind, claudecode and the other coding agents were pretty bad until around Jan of this year (2026). So it's only been about 9 months since devs have had decent coding agents and even less time has elapsed since somewhat wide adoption.


I didn't downvote but I'd expect downvotes for a completely unnuanced thought terminating cliche that ignores everything in the thread. AI fanaticism doesn't help but lots of people don't get downvoted for that alone.

> you will likely put mathematicians out of jobs.

why is that a concern in this context? would you have asked the same about steam engines and horses?

this is a really cool concept, organized very well. and that is very commendable.


steam engines solved a problem people had

If mathematicians aren't solving problems people are having (which your comment seems to imply), then putting them out of their job with AI is not a bad thing. Of course mathematicians are solving problems, just in a very different way than other professions.

Steam engines actually replaced what a horse does with a machine that could do the same job. So far LLMs aren't doing this - they're producing Lean proofs but very little to actually aid in understanding (again so far pretty much all of these proofs have been extremely difficult optimizations of known techniques). Mathematicians do solve a problem that people have: they build theories that explain the world and give us mental models to navigate questions in science, technology, etc., this just isn't a problem that LLMs solve.

So the issue isn't so much that LLMs will replace mathematicians, but that AI companies bragging constantly about how their machines "solve math" will convince people who don't understand the value of math research to no longer fund it, or students who don't yet understand why learning math is useful for developing their brains that it's a waste of time. That could put mathematicians out of a job without providing a useful replacement.

Motto: a mathematician's job isn't to solve the Hodge conjecture, it's to understand why the Hodge conjecture is or isn't true, and turn that understanding into something that makes it easier for the next person to grasp/use/enjoy.

LLMs absolutely have the potential to make this job easier, but the way in which these companies are using them right now risks being antithetical to that goal.


I suppose iphones solved a problem people had as well?

Don't respond to a strawman argument with another strawman. The post you are responding to ignored the reasons given in the second paragraph. They're just trying to score points by preaching to the choir, not engage with the concern.

what strawman?

do you mean,

> All things AI seems to assume that more and faster is better, but there is no justification of that assumption.

is good argument?

of course faster discovery without human in the loop is better. is that not what humans have been optimizing for the past few thousand years ? faster mobility, faster communication, faster medical recovery etc. everything modern civilization has to offer is because of a rush to get better and faster. for example, discovering penicillin 2 years early would've saved ~15 million people more.

why is that not worthy enough to pursue?


i 100% agree with your sentiment. but what is the replacement here? go without any insurance or salary? without providing an alternative for better life, this is basically just robbing someone off a livelihood

If he is 75 and still has to work to get by he has already been robbed of his livelihood. The average lifespan for males in the US is 76ish.

Maybe he just does it for the love of the game, but there are so many old, sick, people in pain who have to labor like this just to get by every day.

Universal basic income is a simple and obvious answer. In a few years we will have zillion dollar companies being run entirely by robots making enormous profits. Share the fucking profits with everyone so that everyone can at least live with dignity.


>Universal basic income is a simple and obvious answer. In a few years we will have zillion dollar companies being run entirely by robots making enormous profits. Share the fucking profits with everyone so that everyone can at least live with dignity.

Sure, but what do we do in the US? UBI would be politically and culturally a non-starter here unless it was extremely means tested and limited (making it no longer universal) and all other forms of welfare were repealed (under the aegis of "reducing government bloat and inefficiency.") Americans won't accept spending their hard earned taxes to make the lives of non-whites, immigrants and the indigent poor too comfortable.


> In a few years we will have zillion dollar companies being run entirely by robots making enormous profits.

"Zillion", I see you've done extensive research here. And these companies aren't making any profit, nevermind enormous.

> Share the fucking profits with everyone so that everyone can at least live with dignity.

Give us a yearly number for what "living with dignity" looks like, multiply that by the entire population, and tell us what you come up with.


How do we prevent UBI from becoming as corrupted and convoluted as taxes?

Governments will use UBI to drive behavior and it will always favor the well connected.


Isn't the whole point of UBI that it is universal and equal, e.g. everyone gets XX$ no matter what? If it is tied somehow to behavior then it is not UBI anymore, but more like a tax credit

He definitely is doing it for the love of the game. Different generation loves to work

tradition are a meme.

It only looks super obvious in hindsight and the well explained blog post. when a team of 5 is tasked with getting a completely new DNS up at the scale and integrate well with cloudflare.

if you spend cycles on nitty gritty opinions like this time to market goes out further and further out. some napkin math, 130 gen13 servers cost "only" ~$2.6M. relative to the importance of the 1.1.1.1 and the market at the time. that is nothing to cloudflare.

this is not to say good system design does not matter. it very much does, but making that call at that time would've butchered the prodcut very much similar to google+, youtube etc.


This one also looks pretty obvious "in foresight" (using the same tools that existed back then. Maybe owner dedupe might be less obvious and require a bit of knowledge and probing into actual data, but for rw vs ro you are fine knowing nothing?) and you forgot the napkin math re. how much your precious "time to market" would have been delayed by.

It's also not nothing, otherwise it would never be optimized away now, but left as is. After all, wasting time on optimization delays "time to market" for other useful features.

I also don't get the reference to YouTube, it's a very successful product, how was it butchered by good system design???


Imagine you're an engineer at cloudflare, an 8 year old (at the time of launch of 1.1.1.1) company. The company is wildly popular and any service launched is going to have a lot of traffic and a lot of attacks right away. Any problems with it are going to embarass the company a lot.

You're tasked with making a DNS caching recursive resolver that can operate at a large scale and will be run on thousands of servers each of which has a lot of GBs of ram.

You are given some period of time to build this and make it production ready. How do you spend your time:

* Focusing on making sure that the resolver works correctly?

* Focusing on make sure that it actually provides improved DNS performance for internet users?

* Handles an very large number of record requests/s?

* Saves a few GB of ram per server?

There are tradeoffs to consider. RAM is cheap, even at today's prices RAM is not the most expensive thing that can go wrong in such a scenario. Having the responses be slow or incorrect is a far more expensive problem. A good engineer would pick a simple data structure that has the right shape but might not be optimal in footprint to focus on correctness and response time. The few extra GBs of RAM per server can be dealt with later.

When building things at scale you want to make sure it works correctly, fails correctly, and does the thing quickly before worrying about reducing resource consumption. I've never seen a project fail on Vec<T> vs Box<[T]> memory differeneces, or even on a few GBs of RAM usage per instance. I have seen them fail on "one wierd corner case of correctness" though, and on poorly thought through failure modes.


> The company is wildly popular and any service launched is going to have a lot of traffic and a lot of attacks right away.

Doesn't this also inform you that your cache will be very large, so you shouldn't use growable structures with slack space when cache entries won't grow; slop space reduces the size of your cache. And also that the query volume will be high so the cached data should require as little work as possible before returning data; spending time marshalling response data on every cache hit increases response time and decreases capacity.


RAM is cheap. I'd find myself far far more concerned with:

* unbounded growth of the cache and properly invalidating after TTL expires (a few GBs of slop is nothing on a server with 64 or more GBs of ram, unbounded growth is a problem).

* making sure the DNS implementation works correctly on both the serving side and recursive resolution side.

* What strategy is best for deduping recursive requests across machines (if something a few miliseconds away has a live result, why do a full lookup taking hundreds or thousands of milliseconds?). This potentially improves RAM usage across the datacenter too from not having a given record on dozens (or more) machines' local cache. I don't know exactly how they do it, but naively I'd look at some sort of DHT shaped solution to look for records in peers within the datacenter. Or maybe some sort of tiered caching with the upper tier being sharded on domain name or the like.

* The biggest performance gains cloudflare can provide in Web and DNS cache come from a cache hit. This is on the order of 10s or 100s of ms due to having a big cache and short distance to the requesting machine. A suboptimal lookup algorithm that is a few microseconds slower in local compute and ram access is just not as important as the other concerns for dedup and cache sharing. That's not to say it's unimportant, just that it's not the top priority when you're trying to deliver this much larger performance gains from other aspects of the system. Thats why they are getting to it several years after release.

Cloudflare writes a lot about distributed systems solutions to various problems. They likely don't think as hard about single machine performance as much as whole datacenter performance when approaching problems.

Keep in mind that the per-server cost of the whole program pre-optimization seems to be about 10GB (from the graph in the post). IME that's not bad for a big busy caching service.


> The biggest performance gains cloudflare can provide in Web and DNS cache come from a cache hit.

Using twice as much ram per cache entry makes the cache half as large, assuming your cache is bounded by ram, unless the queried, unexpired result set is less than the ram budget (which I would tend to doubt... lots of randomized queries out there; maybe I'm wrong if the cache size dropped).

When you're storing billions of records, it makes sense to spend a few minutes to consider how they're used and make a good choice about how to store them.

When you're getting a cache hit tons of times per second, it makes sense to consider every step and which ones don't need to happen every time. You have to consider every step while you're pursing correctness anyway, so might as well have the performance lens active too.

I'm not asking for heroic optimization: I didn't ask for vectorized stuff or kernel/nic offloading or kernel bypass networking... Just you have to use some data structures, you might as well not use ones that are expensive for features you don't need; and you have to store something in your cache, you may as well store something that requires less munging on the way out.

If this were a small local cache, that didn't want to use something already existing like unbound for some reason then yeah, data structures don't make a huge difference, extra marshalling doesn't make a huge difference, just don't reimplement all the CVEs that BIND had in the 90s. But if you're going to allocate 100 TB of ram, make it count. Even if you do use twice the ram but you get value from it, maybe that's fine... I've run wacky systems with bloated storage when there was a benefit. Vec doesn't give any value over a Box<[]> in this case; convenience or lazyness would be fine except that the sheer number of objects makes it worth the few minutes it takes to do something better.


> Using twice as much ram per cache entry makes the cache half as large, assuming your cache is bounded by ram, unless the queried, unexpired result set is less than the ram budget (which I would tend to doubt... lots of randomized queries out there; maybe I'm wrong if the cache size dropped).

This is true. I'm arguing that its unlikely this was ever bound by available RAM. Cloudflare is a DDoS protection company that absorbs attacks. They have a lot of available capacity at any moment. When you're building a service in a sitaution where you have more capacity than you'll likely need.

The savings were 100 TB across >300 data centers. The savings were on the order of 50%. So prior to this reduction the service was using something less than 2/3 of TB per datacenter. The service ram usage was about 10GB per instance according to the graph in post. IDK how cloudflare divides thier stuff between machines, but assuming they don't run less than 64 GB per server that's less than 12 servers per datacenter of ram for a flagship product, and they likely run it spread across 65 of the machines in the datacenter that are also doing other stuff. The per-instance RAM likely isn't the the concerning limit.

Overall RAM usage is proabably a bigger concern. Thats why I would think about dedup between instances and distributed caching strategy first. I could focus on redudcing the ram needed per service instance and get a 50% reduction per machine. Or I could focus on deduping 1/n (where n > 2) reduction in total memory usage across all instances. Personally if I was worried about reducing RAM I'd put more energy into growing N.

However all this is a red herring. The assumption people are making is that the cache was always read-only, and it's obvious that Box<[T]> was the best decision because in a RO cache smaller entries hold more things.

The 1.1.1.1 service advertises improved DNS performace. That's its value add. The biggest performance gain you can have from a cache is not having a cache miss, and in DNS a cache miss means a very expensive recursive lookup. So there's concerns about how to minimize those lookups. If one instance has does a lookup, it makes sense to share that result to the other instances that may need to do a lookup [1]. I don't know off the top of my head if it makes sense to get those updates and modify the existing record or just replace it in the local cache. That comes down to locking strategies and reading patterns in the specific code and service traffic patterns. Until i have hard evidence one way or another I'd like my cache to be able to do both and keep the data structs modifiable until that's nailed down. If per-isntance ram ever becomes the issue, there's easy wins there to buy me time to find better large scale solutions to the problem.

No one is disagreeing that the larger datastructures are larger. No one is disagreeing that they take more RAM, and or even if RAM was the the problem reducing it would be good.

The thing people are pointing out is that this isn't a homework problem about an optimal cache structure in a vacuum. We're pointing our that engineering real large scale solutions has a lot more to consider than a homework problem, and that the thing you're harping about likely didn't have any real budgetary or noticable performance impact on bulding that system. The reduction in ram is just a smallish improvement in operating costs after all the more expensive stuff was figured out.

Put another way 100TB of RAM is ~$350K. Thats one engineer year for a mid-level engineer.[2] Would you rather spend that money to save an equivalent amount of money somewhere, or... would you spend that money putting the engineer on something that saved $700K elsewhere (alternately that generated $700K)?

[1] I talked a lot about dedup and the simple gotcha is "hahah then its not deduped so you need smaller objects". But on a service that is running on a few dozen instances having a few redundant copies to deal with loss of a machine and/or load can still result in 1/(n>2) savings in total ram.

[2] I'm not saying someone worked on this for a year btw, a couple people likely spent a couple months on the code, validation and testing of it. A manager spent time overseeing it. Operations people spent time understaning any effects it had on running systems. Costs add up and it wouldn't suprise me if this didn't end up being roughly break-even for the year.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: