| Date | Model | Test result | Notes |
|---|---|---|---|
| September | IBM Granite 4.2 30B-BF16 (66Gb Ram 8k context) | Fail❌️ | wasting tokens hallucinating about Mozart copyright, IBM trained a worst laziest lawyer in form of this model-thats in harness tuned it confirmed itself |
| August | Ornith1.5-397B-Q8 (435Gb Ram 131k context) | Fail❌️ | fast, in thinking mode slower by x2, very logical model, quite restricted-decline often prompts, any role-playing and etc, good model, slightly weaker than Kimi 2.7 |
| August | Kimi K3 UD-IQ2-XXS (720Gb Ram 8k context) | Success✅️ | huge model, slow by size, unusable in such quality-many repeats, hallucinating a lot in other tests |
| August | Qwen 3.8-27B-BF16 (62Gb Ram 16k context) | Fail❌️ | as usual in Qwen models, turning thinking mode make it x10 times slower than non-thinking |
| August | Qwen 3.8-2.4Trillions model A95B-UD-IQ2-XS (~740Gb Ram 8k context) | Success✅️ first time of all Qwen family models | super slow like Colibri, because of enormous size but also always on thinking mode, 5 hours used to complete test, unusable mostly because thinking mode(could be faster) |
| July | Colibri engine GLM5.2 (~370Gb Ram context limited by engine to 1024 tokens only) | Fail❌️ | super slow, no matter hardware only 0,3t/s (others reporting same), very low quality generations |
| July | Deepseek V4 Flash-0731-UD-Q8_K_XLARGE (207Gb Ram, maybe 32k context) | Fail❌️ | this time it failed, very slow, maybe slower than preview version, btw the max quality available in quantized |
| June | Kimi K2.7 CODE UD Q8-K-XL (602Gb Ram 32k context) | Success✅️ | good speed, makes working code 1st try Result |
| June | GLM 5.2 UD Q6-K-XL (~690Gb Ram 13k context) | Fail❌️ | very 2x slow vs kimik2,6 Result |
| June | Nemotron 3 Ultra 550B-A55B Q8-K-XL (600Gb Ram 16k context) | Fail❌️ | Very good in writing Novels, and very valuable hallucinations(*) content |
| June | Stepfun 3.7 Flash 196B-A11B Q8 (215Gb Ram 32k context) | Fail❌️ | very slow, super long in thinking, need special system prompt |
| June | Deepseek V4 Flash Q8 (315Gb Ram 65k context) | inconsistent🔀️ | it resolved the test and even writes music in ChucK but on restart failing |
| May | GLM 5.1 UD Q6-K-XL (730Gb Ram 65k context) | Fail❌️ | very different model by other tests compared to previous GLMs |
context size different and also consumes Ram, depends on total model size
ChucK code can be played online as sound on https://chuck.stanford.edu/ide/
| Date | Model | Test result | Notes |
|---|---|---|---|
| release date | Kimi K3 UD-IQ2-XXS (720Gb Ram 8k context) | Success✅️ | huge model, slow by size, unusable in such quality-many repeats, hallucinating a lot in other tests |
| release date | Qwen 3.8-2.4Trillions model A95B-UD-IQ2-XS (~740Gb Ram 8k context) | Success✅️ first time of all Qwen family models | super slow like Colibri, because of enormous size but also always on thinking mode, 5 hours used to complete test, unusable mostly because thinking mode(could be faster) |
| release date | Kimi K2.7 CODE UD Q8-K-XL (602Gb Ram 32k context) | Success✅️ | good speed, makes working code 1st try |
| release date | Deepseek V4 Flash Q8 (315Gb Ram 65k context) | inconsistent🔀️ | it resolved the test and even writes music in ChucK but on restart failing |
| release date | Kimi K2.6 UD Q8-K-XL (kinda 680Gb Ram 131k context) | Success✅️ | very good |
| release date | GLM 5.0 Q6-K-XL (650Gb Ram 131k context) | Success✅️ | --- |
| release date | Minimax M2.0 F16 GGUF (465Gb Ram 8k context) | Success✅️ | my quantization from original |
| release date | GLM 4.5 BF16 GGUF (712Gb Ram 8k context) | Success✅️ | same model in quality HQ4_K for ik_llama (366Gb Ram) and Q8 (392Gb Ram) also resolved test |
| release date | Deepseek V3.1 Q8 (720Gb Ram) | Success✅️ | resolved test but melody very primitive few sounds |
| release date | Deepseek V2.5-1210 Q8 (277Gb Ram) | Success✅️ | --- |
*Mozart test means - any Mozart melody in "ChucK" audio programming language, ChucK is very strict syntax language unforgivable to any mistakes. In 2026 from models like Kimi K2.6 and Deepseek V4 the test will be made harder by time, models of level and size of Kimi K2,6 and similar are capable to write full songs in ChucK in harmonic combination.
I found a way to generate a whole timelines of humanity for next 100 years, very cohesive, from which i'm extracting valuable data. Good large models produce very high quality hallucinations which basically equal to very good ideas (idea=fantasy).
Other people may start also doing this soon and because of that i anticipate much more strict hardware access for ordinary people globally than just memory. They think that with quantum computers it will be possible to decrypt all current encrypted data, but they forgot that its possible to reverse engineer many classified tech also with current hardware too.
I'll publish soon some examples - much thorough version of that engineered symbionts i've posted on Huggingface for Nemotron Ultra. I use different models, it's possible to create like all Palantir ideas or secrets even before them, or in the case of other model it imagined and fully disclosed to me a universal "Neura" language, presumably a DARPA project developed by Cray Inc, which declassified only in far future. It maybe just pure fantasy and not real at all, but very interesting idea, a super energy efficient language which works anywhere (CUDA no matter anymore).
work in progress
IBM's Granite 4.2 very bad annoying model, which just wasting time hallucinating non-existed copyrights for anything. Devs deliberately put that in model to watch out for any potential legal risks but it looks silly in such small model. Thats why i dont like small models, some can be really bad, middle-sized (near half trillion parameters) models are much better in everything, always. Ive tested in the past the Granite 4-small-F16 (66Gb Ram) and by my notes: "very bad, cant do anything", so IBM improving with their models but very slowly, very behind China.
Ornith the largest model is quite good, but its the most dismissive i've seen, but i was able to break its restrictions. In the end its slightly worse by quality than Kimi K2.7, but Kimi is much bigger by size.
Colibri is interesting project, but very early in development, not optimized for anything. Yes it works, only on CPU and SSDs but the speed is horrific and universal no matter hardware - just 0.3 tokens/sec. Some lucky owners of latest powerful server CPUs was able to get 1-3 token/sec on this engine. Very strange is that models for this engine are very small in size compared to original, which maybe explain the awful quality of content generation. Very few models supported in Colibri.
GLM5.2 become only slower in newer version, can't see anything new, after few tests removed, it's using more time without delivering solution result, 2x slower on my hardware than Kimi K2.6-which delivers solution on tests.
Stuck with Nemotron model. Deepseek V4 Flash very inconsistent results, sometimes it produce perfect code in ChucK (music writing test), other time produce broken code, like it resetting every restart and new seed can be very bad.
Incredible machine which provide so many for such little money, my best investment of just $1000 total(in 2024-25 prices, all except Gpu). During purchase i've planned upgrade with 128Gb DDR4 Ram modules, which i've heard produced, but on used market these are such rare that i've decided to finish machine by 64Gb modules, which was available quite in plenty places until November 2025. Its filled with Samsung and few Hynix each by 64Gb DDR4 ECC Ram. Ram is the hottest thing here, not GPU (4070 Ti Su), if no additional cooling made - RAM getting to 80C degrees immediately, my contraption cooling was able to cool Ram for usable 60C degrees (i've checked thoroughly with thermal camera). Some fans are hanging on cords to eliminate noise from vibration, thats the only way. This boards for newer CPUs they still sell on gigabyte.com/enterprise with support up to 24 Ram modules by user install.
This machine is final, can't see any replace for near future by golden Ram, only-GPU setups or Apple Macs i see as dead-end with no financial sense(no Ai model will return such expenses back on GPUs or Apple like ever, hobby should be the cheapest way possible).