Measuring What Matters: Towards a Common Evaluation Framework for Christian LLMs
Existing AI benchmarks don't focus on Christian faith and human flourishing, making it hard to measure how well large language models handle biblical knowledge, theology, and pastoral sensitivity. This talk, co-presented with Nick Skytland (Gloo), proposes an open, collaborative framework for evaluating Christian LLMs: a set of roughly 3,000 questions across seven human flourishing dimensions, together with tooling to automate evaluation using a diverse panel of LLM judges and a defined rubric. It reports early progress testing leading models (including Gemini, DeepSeek, Grok, and Mistral) and invites others to help define shared standards, contribute data, and build this evaluation framework together.