Friday and Saturday I was at Socrates Austria. The claim floating around the room was that Jev (TypeSafe AI) can play Doom – not through pixels, but through structured game state (enemy positions, distances, angles as JSON) that Jev translates into an action ~10x/second. Impressive stuff. Someone in the crowd mentioned Laya: supposedly the open-source answer to Jev, same model category, “System 1” is what they call it, typed values instead of free text. Tried it on the spot on a laptop, just to see if Laya could pull off something comparable. Failed pretty badly. No clean setup, no measurement, just the impression: this thing doesn’t react to the state at all. Still in the session, a second, equally quick reality check on something more mundane: classifying a support ticket. Similarly underwhelming. No clear differentiation, flat answers, basically the same output no matter what you feed it. So I wanted to know for real, properly set up instead of improvised. No Doom, I don’t have a ViZDoom setup lying around, but Flappy Bird as an obviously simpler test with the same core idea: structured state instead of pixels, Laya interprets it and outputs the action (flap or don’t). That’s the experiment documented below. Spoiler: the numbers at the end look a lot better than the improvised first attempt — but the underlying skepticism from that weekend only partially went away.
Laya isn’t a classic LLM, it’s a non-autoregressive, bidirectional encoder (backbone: answerdotai/ModernBERT-large, ~400M parameters) that returns one of three answer types for a given state – choice (selection + probabilities), score (ordinal rating), or noul (calibrated yes/no probability). According to the docs it’s built for email routing, spam guardrails, support ticketing – classic business classification, no vision, no control tasks. Exactly the category where the quick support-ticket test had already made me suspicious. Read more




