As artificial intelligence becomes more capable and widely used, there is growing interest in ensuring that AI systems behave safely and consistently with the values their developers publicly commit to. AI Character Evaluations (AICE) will independently evaluate whether leading AI systems meet those commitments, helping improve transparency and public trust in advanced AI.
Improving accountability for AI systems
Many leading AI companies publish documents that explain how their AI models are intended to behave. These documents, known as “Model Specs”, set out the principles, behaviours and safety standards developers expect their systems to follow.
For example, OpenAI’s Model Spec outlines that its models should avoid facilitating serious harm and should comply with applicable laws, while Anthropic’s Constitution sets out the values and ethical principles intended to guide Claude’s behaviour.
AICE will develop methods to independently assess whether AI systems are following these published standards. It will publish reports and research that help developers improve model adherence to their stated principles, increase transparency around the safety of frontier AI systems, and contribute research on how advanced AI should behave in high-stakes situations.
Robert McCarthy, founder of AICE, said: “As AI systems become increasingly powerful and widely deployed, their character and behaviour will shape the world we live in. At AI Character Evaluations, our mission is to encourage the development of model specifications and AI behaviours that benefit society and help reduce societal-scale risks from advanced AI.”
Building on AI safety research at UCL
Robert’s PhD research at UCL has focused on improving the safety monitoring and evaluation of frontier large language models.
His work on chain-of-thought monitoring was supported by a £146,000 grant from the UK AISI Challenge Fund. In collaboration with OpenAI, he demonstrated that reasoning training can improve the monitoring of AI systems by reducing their ability to actively control or conceal their reasoning processes.
His research examining whether large language models can encode hidden communications and reasoning has been published at NeurIPS 2025 and received the Social Impact Award at IJCNLP-AACL 2025.
Dr Zhibin (Alex) Li, Associate Professor at UCL and Robert’s PhD supervisor, said: “Robert’s success in securing this funding is strong evidence of the important AI safety research taking place at UCL. Through his PhD work, he has built deep expertise in evaluating frontier large language models, which he will take into his new organisation. I’m confident AICE’s independent evaluations will deliver real public benefit by helping ensure that frontier AI systems are safe and trustworthy.”
Further Reading
Open AI model spec
Anthropic - Claude’s Constitution
UK AISI Challenge Fund
Open AI Reasoning Models chain of thought controllability
NeurIPS 2025 paper: Large language models can learn and generalize steganographic chain-of-thought under process supervision
IJCNLP-AAACL 2025 Best Papers