Public
Work
9 public workstreams. Dated, named, and scored when we can score them.
2026-08Sudoku SFT + RL EnvironmentA verifiable Sudoku environment for supervised fine-tuning and reinforcement learning. Constraints you can score, not vibes.Environments2026-07LangSmith Evaluation SystemHow Solerno measures agents in production: traces, datasets, judges, and a loop back into the next run.Evals & safety2026-06DeepTeam Red-Teaming with Local QwenAdversarial testing against a local uncensored Qwen 3.8 model. Safety work that does not depend on a hosted API’s filters.Evals & safety2026-05Agent OntologyA shared vocabulary for agents, tools, memory, and environments, so systems can compose instead of colliding.Environments2026-05Pi + Orca Local EnvironmentThe lab’s local development environment: agent orchestration, worktrees, and a machine you can actually see.Physical & local2026-04FigmaBenchA 100-task benchmark for whether language models can design in Figma through the Plugin API. Not screenshot-to-code.Environments2026-04SO-101 Robot ArmA tabletop arm on NVIDIA Orin Nano 8GB, ROS 2, and Gemini Flash. Physical intelligence at lab scale, not a robotics company claim.Physical & local2026-03İnanç Tekstil Curtain StoreA live made-to-measure curtain shop: pick the fabric, set width and pleat, then try it in a photo of your own room before you buy.Applied2026-02Agentic Due DiligenceLong-running analysis over heterogeneous data sources. Agents that read, wait, and write a file, not a chatbot with a search bar.Applied