If you go to England, you're better off going to the New Forest in Hampshire. Despite the name, it was established during the Norman period and at least feels like a substantial tract of older woodland. (Tiny by North American standards of course.)
I am really surprised you’re seeing half the generation performance of 3.6 with 3.8 at the same parameter size and quantization (and same prefill performance to boot) - is there just an optimization in the stack somewhere for 3.6 that hasn’t landed yet for 3.8?
Curious if other folks are also surprised by this apparent discrepancy or have a ready explanation.
I suspect this might be due to MTP mispedictions. Either because the 3.8 quantized model weights do not include MTP heads, or as what happened with Ornith-1.5 recently, corrupted MTP heads, or due to a software issue in Ollama.
I remember an article a couple weeks back where someone used an HPE server with two older Xeon E-series CPUs and got reasonable performance by compiling llama with optimal switches for the architecture (memory alignment, page sizes, etc). I was impressed because those Xeons only had AVX2. If you are willing to spend a little more, you can get a slightly newer Xeon with AVX-512 and 4 sockets.
I can't find the post though.
Edit: a quick search from a server refurbisher nearby gives me a dual Xeon Gold 6330 28-core machine with 512GB and a 16GB V100 GPU. The memory is spread across 32 slots, for about £6,470.
If I understand correctly, it’s not that the LLM can’t write a good comment, it’s that you want to be able to interpret and understand the generated code without comments - and in that process end up writing comments yourself.
I don’t get that argument. Most of the time by the end of the session the comments from the agent encode tricky details that I told the agent to write down so it stops making “simplifying” assumptions. Thus comments at the end of a couple days of agent-only coding, when I start to actually read and edit the prose, contain the details which aren’t possible to know from reading the local code. It may help that my last couple sessions before I start reading the code myself are variations on telling the agent to self-review and improve the comments in specific ways, so the comments left are only those the agent thought remain meaningful at clarifying unexpected interactions between the local code and other code that needs to be referenced to understand it.
Core Series 3 processors, Thunderbolt 4, a new keyboard with fingerprint reader and backlight.
As a current Framework 12 user I'm not in a huge rush to upgrade the mainboard, but will be snagging the backlit keyboard as soon as it's available by itself. Would love to see a haptic trackpad one day, as well as a screen that covers a larger color space.
>Ansel has not published a stable release yet and there is no ETA for one : a new release is published when the list of all bugs have been cleared, so we know the software and stable. So far, Ansel only publishes revisions, which are intermediate states of the sourcecode.
I have a similar question and I’m inferring the answer is no - look at the cache hit rate of 23% for the 128GB M5 Max. I had previously assumed that the 40B active meant that a set of layers was chosen as THE expert for a given prompt and generation was then limited to those layers until complete. But in that case you’d have expected the expert caching to have a super high hit rate once you had enough RAM to hold an entire expert’s worth of layers.
reply