Discussion about this post

User's avatar
Brian Heming's avatar

These are incredibly lobotomized models, with poor performance due to how much attention was spent lobotomizing them (i.e. "safety") versus making them not suck.

Check it out: the 120 billion parameter model performs way worse on creative writing than Google's 4 billion parameter model. Blech. https://eqbench.com/creative_writing.html ( search for gpt-oss-120b and gemma-3-4b-it )

Nibmeister's avatar

Cool. Perhaps the 120B model could be run and retrained on NVidia's DGX Spark, if this little AI box ever gets released.

8 more comments...

No posts

Ready for more?