Environments2026-04
FigmaBench
A 100-task benchmark for whether language models can design in Figma through the Plugin API. Not screenshot-to-code.
FigmaBench scores generated designs across six layers: element matching, layout, design tokens, schema similarity, screenshot comparison, and a VLM judge. L1–L4 run from JSON alone, so they can train offline. L5–L6 need a live Figma instance.
The public site is an interactive explainer: a golden-versus-generated replay, a stacked leaderboard, a task browser, and playable scoring widgets. It is the first Solerno research surface that already feels like a lab, not a brochure.
This is the flagship. New work should be this concrete.