FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025,
https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions]
https://github.com/SalesforceAIResearch/FaithEval