Is your LLM really clever? Can it mark its own homework?
ELOQUENT Lab 2027
4th edition of the lab for evaluation of generative language model quality at CLEF, the Conference and Labs of the Evaluation Forum
Here's how to participate
Sign up here!

Participating in the tasks can be done as a simple one-off experiment on one or more of the tasks or a more elaborate experimental approach, depending on how you want to work with the challenges.

Experimental reports will be published in the working notes of the workshop, but there is no requirement to submit a formal report; if experiments involve hypothesis testing and exploration of more lasting value, they can be revised and later published elsewhere in more archival channels.

CLEF Logo

Task Voight-Kampff

Can your LLM fool a classifier to believe it is human?

This task explores whether automatically-generated text can be distinguished from human-authored text, and is organised in collaboration with the PAN lab at CLEF.

  • More information about the task will be released as details for this year's edition have been settled.
  • part human, part machine

    Robustness Task Cultural Robustness and Diversity

    Will your machine respond with the same content to all of us?

    This task has run in three variants in previous editions, and this year will tests how well a model achieves consistency across several languages and how well the model adapts to the local culture of a linguistic area.

  • More information about the task will be released as details for this year's edition have been settled.
  • janus, a two-faced deity, depicted on a roman coin

    Sensemaking Task Generating and Scoring Exams

    Can your language model prep, sit, or rate an exam for you?


    Pavel Šindelář, Markarit Vartamperian

    • More information about the task will be released as details for this year's edition have been settled.

    study session

    Student projects

    Do you teach a class related to generative language models? Do you supervise students interested in generative language models? Are you a student searching for a project?

    a teacher

    Some of the ELOQUENT tasks are suitable for use as a class assignment or as a diploma project. Get in touch with us for suggestions of extensions and other ideas!

    Voight-Kampff, Hallucigen, Robustness and Consistency, and Topical Quiz tasks

    Voight-Kampff, Robustness and Consistency, Preference Prediction, and Sensemaking tasks

    Voight-Kampff, Cultural Robustness and Consistency, PISA Topical Quiz generation, and Sensemaking tasks

    Organising Committee
    some people at a work table
    • AMD Silo AI: Maria Barrett, Jussi Karlgren, Aarne Talman, Georgios Stampoulidis
    • Charles University: Ondřej Bojar, Věra Kloudová, Jindřich Libovický, Andrei Manea, Lucie Poláková, Pavel Šindelář, Gianluca Vico
    • University of Amsterdam: Bruno Sotić
    • Université Grenoble Alpes: Diandra Fabre, Lorraine Goeuriot, Philippe Mulhem, Didier Schwab, Markarit Vartampetian
    • University of Ljubljana and Jožef Stefan Institute: Nikola Ljubešić, Špela Vintar
    • University of Oslo: Jindřich Helcl, Yves Scherrer, Erik Velldal, Lilja Övrelid
    • University of Tartu: Mark Fishel
    • University of Tennessee: Rohit Raj Gunti
    • Université de Toulouse: Josiane Mothe

    Contact us at eloquent-clef2027-organizers AT googlegroups.com

    Thank you

    The ELOQUENT lab is partially supported by the OpenEuroLLM and the DeployAI projects through their activities on building, evaluating, and disseminating generative language models.

    Page layout from Codepen.