Angelina Parfenova

I study how large language models reason, communicate, and make decisions in settings involving human interpretation. My work develops evaluation methods that go beyond benchmark accuracy to examine reliability, diversity of reasoning, qualitative analysis, and human-AI alignment. By combining NLP with ideas from computational social science and cognitive science, I aim to build methods that help us understand not only whether language models work, but how and when they can be trusted in human-centered applications.