Solerno
Applied AI labAmsterdam, NL
Systems you can score.
Solerno builds AI systems you can evaluate: environments, agents, and machines.
LatestSudoku SFT + RL Environment2026-08
FigmaBenchSudoku SFT + RL EnvironmentLangSmith Evaluation SystemDeepTeam Red-Teaming with Local QwenAgent OntologyPi + Orca Local EnvironmentSO-101 Robot Armİnanç Tekstil Curtain StoreAgentic Due DiligenceFigmaBenchSudoku SFT + RL EnvironmentLangSmith Evaluation SystemDeepTeam Red-Teaming with Local QwenAgent OntologyPi + Orca Local EnvironmentSO-101 Robot Armİnanç Tekstil Curtain StoreAgentic Due Diligence
Live
Featured
FigmaBench
A 100-task benchmark for whether language models can design in Figma through the Plugin API. Not screenshot-to-code.
- tasks
- 100
- reward layers
- 6
- UI categories
- 10
More work
Work
Named systems, not services. Each entry is something we trained, measured, or shipped.
2026-08Sudoku SFT + RL EnvironmentA verifiable Sudoku environment for supervised fine-tuning and reinforcement learning. Constraints you can score, not vibes.Environments2026-07LangSmith Evaluation SystemHow Solerno measures agents in production: traces, datasets, judges, and a loop back into the next run.Evals & safety2026-06DeepTeam Red-Teaming with Local QwenAdversarial testing against a local uncensored Qwen 3.8 model. Safety work that does not depend on a hosted API’s filters.Evals & safety2026-05Agent OntologyA shared vocabulary for agents, tools, memory, and environments, so systems can compose instead of colliding.Environments2026-05Pi + Orca Local EnvironmentThe lab’s local development environment: agent orchestration, worktrees, and a machine you can actually see.Physical & local2026-04SO-101 Robot ArmA tabletop arm on NVIDIA Orin Nano 8GB, ROS 2, and Gemini Flash. Physical intelligence at lab scale, not a robotics company claim.Physical & local2026-03İnanç Tekstil Curtain StoreA live made-to-measure curtain shop: pick the fabric, set width and pleat, then try it in a photo of your own room before you buy.Applied2026-02Agentic Due DiligenceLong-running analysis over heterogeneous data sources. Agents that read, wait, and write a file, not a chatbot with a search bar.Applied