Rendered at 09:40:54 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
arikrahman 6 hours ago [-]
I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.
mmastrac 3 hours ago [-]
I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.
TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems
TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems