Frequently Asked Questions

Task 1 & 2

Q1. Which LLM and RadFact implementation are used for evaluation?

A1.The evaluation pipeline provided in the shared evaluation code currently uses GPT-4.1 mini as the RadFact entailment judge. This was the organizers’ original model choice.

The organizers are also evaluating GPT-5.6-sol with low reasoning effort to determine whether it provides more accurate results for this specific task. Any change to the final evaluation setup will be communicated to participants.

Participants who wish to align their local validation with the official setup should use the RadFact pipeline included in the shared evaluation code.

Q1. Which ground-truth report is used for the hidden test set?

The ground truth is taken from the English reports. For each patient in the hidden test set, only one report is available and used as the reference report.