where do reasoning models refuse? Investigating where safety decisions are made in reasoning models, through statistical and mechanistic techniques, uncovering interesting differences in reasoning patterns! fuzzing harness generator for patch completeness Automatically generating a set of fuzzing harnesses conditioned on the root cause of a vulnerability to surface sibling bugs in overfitting patches. project canary Collecting TTPs and behavioural fingerprints from a bunch of AI agents interfacing with vulnerable web apps, and mapping them onto the MITRE ATT&CK matrix playbooks for AI red and blue teaming Summarising the playbooks for AI red and blue teaming developed by my team at The Alan Turing Institute and collaborators from The MITRE Corporation. film I've taken my film camera on a few trips now, here's some that I've taken that I'm particularly fond of.