About me

Anyone who wants to know who is behind CrucibleMark and why a UX designer is benchmarking language models should read on.

Since 2003 I have been working as a freelancer at Media-Garage in the areas of UX design, web design, and print design. A consistent focus has been the visualization of facts, data, and texts — from books and magazines to data sheets and complex websites that I developed together with clients in terms of both content and design.

What has always interested me is the connection between structure, typography, and web typography. Later I moved deeper into the UX world, driven in part by a market in which low-cost providers and quick solutions were becoming ever louder. That is precisely why considered, well-thought-out design with substance has remained important to me.

Through the open-source CMS Drupal I came into intensive contact with the open-source world and encountered a community whose willingness to help and technical energy still impresses me today. That attitude — open, pragmatic, collaborative — fits well with the way I think about digital tools and systems.

Professionally, I now work in an environment where I help shape frontends, content, and digital access for a broad community. That gives me the good feeling of working on something useful that is genuinely needed in everyday life.

In parallel, over the past few months I have experienced how significantly AI assistants have changed my work. Especially as a solo practitioner in the frontend space, they help me become more productive, clear technical hurdles faster, and refocus on what I do best: interfaces, type systems, grids, structure, and clarity.

CrucibleMark also emerged from this experience. For me, the project is not just a benchmark but an experiment: how far can I get with AI as an assistance system, and where do local open-weight models actually stand in that picture? As an amplifier of my strengths and a counterweight to my weaknesses. I wanted to know which models genuinely help me in everyday work, how reliable they are, and what they actually deliver compared to commercial offerings.

Kay Beißert, author

Addendum: why do local models perform better than expected?

When I started evaluating models for my own everyday use, one thing about classic LLM leaderboards bothered me: they implicitly answer the question "Which model is best overall?" — but that was never my actual question. What I wanted to know was: is this model, which I can host locally on my own hardware, good enough for exactly the task I carry out every day?

An aggregate score across seven to nine fundamentally different disciplines structurally favors generalists — typically large, expensive frontier models with correspondingly large training budgets. A small, specialized model on my own GPU never had a fair chance in that comparison, even if it was on par with or even better than a frontier model in the one discipline that matters to me — code reviews, CLI automation, documentation.

That is why the coding, developer, and agentic profiles are not an add-on feature for me but the actual core of CrucibleMark: they do not answer "Which model is best?" but rather "Is this model good enough for my task?" — the question that matters when I am weighing whether to self-host a model or pay for an API.