Aakib Ansari.

Hands-on Tests

Every article here is a real test — we run the model ourselves, show the verbatim prompt, the unedited output, and give an honest assessment of what worked and what didn't. No cherry-picking, no cleanup. These are the tests that show you what a model can actually do, not what the press release says it can do.

Same Prompt, Two Models: Gemini 3.1 Pro vs GLM-5.2 Both One-Shot a Tower Defense Game — With Opposite Architectures
Hands-on Test9 min read
Same Prompt, Two Models: Gemini 3.1 Pro vs GLM-5.2 Both One-Shot a Tower Defense Game — With Opposite Architectures

We gave Gemini 3.1 Pro and GLM-5.2 the exact same tower defense prompt, both at max/high effort. Both one-shot a fully playable game with zero follow-up fixes — but they made opposite architectural choices on the one open question the prompt left them, and GLM-5.2 quietly added a fifth enemy type and a damage-type counter system nobody asked for.