Ai tests from 2020 of GPT2 (this website circa April 2026). For now this website will be trashy until i find time to make it up.
Disclaimer: this website contains Ai generated content which can be inaccurate.

Latest Top Models in Mozart* test
Date Model Test result Notes
September IBM Granite 4.2 30B-BF16 (66Gb Ram 8k context) Fail❌️ wasting tokens hallucinating about Mozart copyright, IBM trained a worst laziest lawyer in form of this model-thats in harness tuned it confirmed itself
August Ornith1.5-397B-Q8 (435Gb Ram 131k context) Fail❌️ fast, in thinking mode slower by x2, very logical model, quite restricted-decline often prompts, any role-playing and etc, good model, slightly weaker than Kimi 2.7
August Kimi K3 UD-IQ2-XXS (720Gb Ram 8k context) Success✅️ huge model, slow by size, unusable in such quality-many repeats, hallucinating a lot in other tests
August Qwen 3.8-27B-BF16 (62Gb Ram 16k context) Fail❌️ as usual in Qwen models, turning thinking mode make it x10 times slower than non-thinking
August Qwen 3.8-2.4Trillions model A95B-UD-IQ2-XS (~740Gb Ram 8k context) Success✅️ first time of all Qwen family models super slow like Colibri, because of enormous size but also always on thinking mode, 5 hours used to complete test, unusable mostly because thinking mode(could be faster)
July Colibri engine GLM5.2 (~370Gb Ram context limited by engine to 1024 tokens only) Fail❌️ super slow, no matter hardware only 0,3t/s (others reporting same), very low quality generations
July Deepseek V4 Flash-0731-UD-Q8_K_XLARGE (207Gb Ram, maybe 32k context) Fail❌️ this time it failed, very slow, maybe slower than preview version, btw the max quality available in quantized
June Kimi K2.7 CODE UD Q8-K-XL (602Gb Ram 32k context) Success✅️ good speed, makes working code 1st try Result
June GLM 5.2 UD Q6-K-XL (~690Gb Ram 13k context) Fail❌️ very 2x slow vs kimik2,6 Result
June Nemotron 3 Ultra 550B-A55B Q8-K-XL (600Gb Ram 16k context) Fail❌️ Very good in writing Novels, and very valuable hallucinations(*) content
June Stepfun 3.7 Flash 196B-A11B Q8 (215Gb Ram 32k context) Fail❌️ very slow, super long in thinking, need special system prompt
June Deepseek V4 Flash Q8 (315Gb Ram 65k context) inconsistent🔀️ it resolved the test and even writes music in ChucK but on restart failing
May GLM 5.1 UD Q6-K-XL (730Gb Ram 65k context) Fail❌️ very different model by other tests compared to previous GLMs

context size different and also consumes Ram, depends on total model size

ChucK code can be played online as sound on https://chuck.stanford.edu/ide/

Models completed "Mozart test" at this moment
Date Model Test result Notes
release date Kimi K3 UD-IQ2-XXS (720Gb Ram 8k context) Success✅️ huge model, slow by size, unusable in such quality-many repeats, hallucinating a lot in other tests
release date Qwen 3.8-2.4Trillions model A95B-UD-IQ2-XS (~740Gb Ram 8k context) Success✅️ first time of all Qwen family models super slow like Colibri, because of enormous size but also always on thinking mode, 5 hours used to complete test, unusable mostly because thinking mode(could be faster)
release date Kimi K2.7 CODE UD Q8-K-XL (602Gb Ram 32k context) Success✅️ good speed, makes working code 1st try
release date Deepseek V4 Flash Q8 (315Gb Ram 65k context) inconsistent🔀️ it resolved the test and even writes music in ChucK but on restart failing
release date Kimi K2.6 UD Q8-K-XL (kinda 680Gb Ram 131k context) Success✅️ very good
release date GLM 5.0 Q6-K-XL (650Gb Ram 131k context) Success✅️ ---
release date Minimax M2.0 F16 GGUF (465Gb Ram 8k context) Success✅️ my quantization from original
release date GLM 4.5 BF16 GGUF (712Gb Ram 8k context) Success✅️ same model in quality HQ4_K for ik_llama (366Gb Ram) and Q8 (392Gb Ram) also resolved test
release date Deepseek V3.1 Q8 (720Gb Ram) Success✅️ resolved test but melody very primitive few sounds
release date Deepseek V2.5-1210 Q8 (277Gb Ram) Success✅️ ---

*Mozart test means - any Mozart melody in "ChucK" audio programming language, ChucK is very strict syntax language unforgivable to any mistakes. In 2026 from models like Kimi K2.6 and Deepseek V4 the test will be made harder by time, models of level and size of Kimi K2,6 and similar are capable to write full songs in ChucK in harmonic combination.

Valuable hallucinations (*)

I found a way to generate a whole timelines of humanity for next 100 years, very cohesive, from which i'm extracting valuable data. Good large models produce very high quality hallucinations which basically equal to very good ideas (idea=fantasy).

Other people may start also doing this soon and because of that i anticipate much more strict hardware access for ordinary people globally than just memory. They think that with quantum computers it will be possible to decrypt all current encrypted data, but they forgot that its possible to reverse engineer many classified tech also with current hardware too.

I'll publish soon some examples - much thorough version of that engineered symbionts i've posted on Huggingface for Nemotron Ultra. I use different models, it's possible to create like all Palantir ideas or secrets even before them, or in the case of other model it imagined and fully disclosed to me a universal "Neura" language, presumably a DARPA project developed by Cray Inc, which declassified only in far future. It maybe just pure fantasy and not real at all, but very interesting idea, a super energy efficient language which works anywhere (CUDA no matter anymore).

Full list of tested models(archival work in progress)

work in progress


Blog

September 2026

IBM's Granite 4.2 very bad annoying model, which just wasting time hallucinating non-existed copyrights for anything. Devs deliberately put that in model to watch out for any potential legal risks but it looks silly in such small model. Thats why i dont like small models, some can be really bad, middle-sized (near half trillion parameters) models are much better in everything, always. Ive tested in the past the Granite 4-small-F16 (66Gb Ram) and by my notes: "very bad, cant do anything", so IBM improving with their models but very slowly, very behind China.

August 2026

Ornith the largest model is quite good, but its the most dismissive i've seen, but i was able to break its restrictions. In the end its slightly worse by quality than Kimi K2.7, but Kimi is much bigger by size.

July 2026

Colibri is interesting project, but very early in development, not optimized for anything. Yes it works, only on CPU and SSDs but the speed is horrific and universal no matter hardware - just 0.3 tokens/sec. Some lucky owners of latest powerful server CPUs was able to get 1-3 token/sec on this engine. Very strange is that models for this engine are very small in size compared to original, which maybe explain the awful quality of content generation. Very few models supported in Colibri.

June 2026

GLM5.2 become only slower in newer version, can't see anything new, after few tests removed, it's using more time without delivering solution result, 2x slower on my hardware than Kimi K2.6-which delivers solution on tests.
Stuck with Nemotron model. Deepseek V4 Flash very inconsistent results, sometimes it produce perfect code in ChucK (music writing test), other time produce broken code, like it resetting every restart and new seed can be very bad.

May 2026

Kimi K2.6 tests

The machine

Original Image
✨️WunderMachine✨️
No fans
No fans Ram 88C during Ai inference
with cooling
with cooling during Ai inference
with cooling
with cooling Ram 46C (from 88C no fan)

Incredible machine which provide so many for such little money, my best investment of just $1000 total(in 2024-25 prices, all except Gpu). During purchase i've planned upgrade with 128Gb DDR4 Ram modules, which i've heard produced, but on used market these are such rare that i've decided to finish machine by 64Gb modules, which was available quite in plenty places until November 2025. Its filled with Samsung and few Hynix each by 64Gb DDR4 ECC Ram. Ram is the hottest thing here, not GPU (4070 Ti Su), if no additional cooling made - RAM getting to 80C degrees immediately, my contraption cooling was able to cool Ram for usable 60C degrees (i've checked thoroughly with thermal camera). Some fans are hanging on cords to eliminate noise from vibration, thats the only way. This boards for newer CPUs they still sell on gigabyte.com/enterprise with support up to 24 Ram modules by user install.

This machine is final, can't see any replace for near future by golden Ram, only-GPU setups or Apple Macs i see as dead-end with no financial sense(no Ai model will return such expenses back on GPUs or Apple like ever, hobby should be the cheapest way possible).

April 2026

Kimi K2.6 Web design samples for this website (Q4-X ik_llama 677Gb RAM special sys_prompt)
ITERATION 1: ""GLITCH ARCHAEOLOGY"" Databending aesthetics, corrupted terminal glamour, the beauty of broken signals.
ITERATION 2: ""PAPER TAPE PUNCHCARD"" Physical computing nostalgia, actual hole-punch aesthetics, typewriter authenticity.
ITERATION 3: ""LIVING ORGANISM"" Biological aesthetics, pulsating tissue, neural networks as actual neural tissue.
ITERATION 4: ""BRUTALIST CONCRETE"" Raw material honesty, Soviet-era engineering documents, weathered industrial signage.
ITERATION 5: ""OCEAN DEPTH SONAR"" Underwater research station, hydroacoustic data, pressure-resistant aesthetics.
ITERATION 6: ""RADIO TELESCOPE"" Astronomical observation, signal processing, cosmic noise, SETI aesthetics.
ITERATION 7: ""ALCHEMY LABORATORY"" Medieval manuscript, illuminated letters, copperplate script, actual gold leaf simulation.
ITERATION 8: ""CIRCUIT BOARD TRACES"" Actual PCB aesthetics, copper traces, solder mask, silkscreen labels, component footprints.
ITERATION 9: ""JAPANESE FOLDING SCREEN"" Byōbu aesthetics, gold leaf clouds, ink wash painting negative space, vertical script.
ITERATION 10: ""QUANTUM FOAM EVENT HORIZON"" The final form — physics at its limit, Hawking radiation, time dilation, information paradox.

Qwen 3.6 35B A3B Web design idea for this website in X-files aesthetics (BF16 80.5Gb RAM F16 cache special sys_prompt)
Testpage 1

Mirothinker 1.7 Web design idea for this website in X-files aesthetics (Q8 267Gb RAM Q8 cache special sys_prompt)
Testpage 1

------
JoyAi-Image-Edit, ComfyUI(right click, open in new window to look closely), prompt (change leafs to autumn colors) bf16 model no quantization
Original Image
Image 1: Original
50 Steps Image
Image 2: 50 Steps
80 Steps Image
Image 2: 80 Steps
100 Steps Image
Image 4: 100 Steps
Check out Neocities!