I am Ram Bharadwaj, a software engineer turned AI safety researcher. My current work focuses on understanding how metagaming emerges in large language models, and on developing methods to automatically search for misalignments.
My publications are available on Google Scholar here, and my CV is available here. You can also find me on LinkedIn and Twitter.