About me
Anyone who wants to know who is behind CrucibleMark and why a UX designer is benchmarking language models should read on.
Since 2003 I have been working as a freelancer at Media-Garage in the areas of UX design, web design, and print design. A consistent focus has been the visualization of facts, data, and texts — from books and magazines to data sheets and complex websites that I developed together with clients in terms of both content and design.
What has always interested me is the connection between structure, typography, and web typography. Later I moved deeper into the UX world, driven in part by a market in which low-cost providers and quick solutions were becoming ever louder. That is precisely why considered, well-thought-out design with substance has remained important to me.
Through the open-source CMS Drupal I came into intensive contact with the open-source world and encountered a community whose willingness to help and technical energy still impresses me today. That attitude — open, pragmatic, collaborative — fits well with the way I think about digital tools and systems.
Professionally, I now work in an environment where I help shape frontends, content, and digital access for a broad community. That gives me the good feeling of working on something useful that is genuinely needed in everyday life.
In parallel, over the past few months I have experienced how significantly AI assistants have changed my work. Especially as a solo practitioner in the frontend space, they help me become more productive, clear technical hurdles faster, and refocus on what I do best: interfaces, type systems, grids, structure, and clarity.
CrucibleMark also emerged from this experience. For me, the project is not just a benchmark but an experiment: how far can I get with AI as an assistance system, and where do local open-weight models actually stand in that picture? As an amplifier of my strengths and a counterweight to my weaknesses. I wanted to know which models genuinely help me in everyday work, how reliable they are, and what they actually deliver compared to commercial offerings.
Addendum: why do local models perform better than expected?
When I started evaluating models for my own everyday use, one thing about classic LLM leaderboards bothered me: they implicitly answer the question "Which model is best overall?" — but that was never my actual question. What I wanted to know was: is this model, which I can host locally on my own hardware, good enough for exactly the task I carry out every day?
An aggregate score across seven to nine fundamentally different disciplines structurally favors generalists — typically large, expensive frontier models with correspondingly large training budgets. A small, specialized model on my own GPU never had a fair chance in that comparison, even if it was on par with or even better than a frontier model in the one discipline that matters to me — code reviews, CLI automation, documentation.
That is why the coding, developer, and agentic profiles are not an add-on feature for me but the actual core of CrucibleMark: they do not answer "Which model is best?" but rather "Is this model good enough for my task?" — the question that matters when I am weighing whether to self-host a model or pay for an API.