Bruno Wassermann

Senior Software Engineer · IBM Israel Research Lab

I work on making large systems behave. First the distributed infrastructure, then the monitoring around it, and now the cost and reliability of AI agents.

Right now I lead a small team working on reducing the cost of agentic workflows without sacrificing task quality. We are developing an uncertainty-aware router that decides at each step whether a smaller open-weight model is likely to be sufficient or whether the task should be escalated to a larger model. The routing logic itself is only part of the problem. Much of the work is in designing controlled experiments that measure the savings and the effect on task success reliably. We are also working on integrating this into llm-d, with the aim of making it a practical capability for people running open-weight models rather than leaving it as an experimental result.

Before that, I spent roughly a decade working on cloud systems. I helped build Watson Developer Cloud from its first open beta through its subsequent growth, and worked with a team of two to four engineers to operate its core distributed services with close to zero downtime. That experience shaped how I think about operational systems: monitoring should help people identify problems, not page them whenever a metric moves.

I later moved into AIOps, working on anomaly detection and root-cause analysis over logs and metrics. My team put this into production together with the group operating IBM Cloud's network. I also built Clue, a release-verification system that detects when a change to a machine training pipeline degrades customers' chatbot models before the change is released.

More recently, I led LakeGuard, a project on access control for structured and vector data in GenAI lakehouses. We took permissions originating in systems such as Box and SharePoint, mapped them into a common model, and enforced them across Iceberg, Milvus, Presto and Postgres while exploring how to carefully manage query latency.

I have a PhD from UCL on data-driven detection and diagnosis of failures in service compositions, six granted patents, and papers at EuroSys, ICSE, WWW and ICSOC. Earlier in my career, I was an Eclipse committer on the BPEL Designer, worked on Grid workflow tooling at UCL, and started out at Goldman Sachs, where I learned how much a production release can actually cost.

The common thread in my work is curiosity about difficult technical problems and a preference for problems that matter in practice. I am most at home when a problem is still poorly defined: turning it into questions that can be tested, working through the evidence, and then building something that holds up in practice. I like small teams where people trust and support one another, feel safe disagreeing openly, and keep getting better at designing experiments that tell us early when we are wrong.