CQ4OE Benchmark

A benchmark for assessing LLM-assisted ontology generation from competency questions

A reproducible benchmark for evaluating LLM-generated ontologies against CQ-aligned gold standards

Six source ontologies with explicit provenance from each competency question to the classes, properties, and axioms required to answer it. Two evaluation tasks. CQ2Term measures CQ-specific term prediction over 99 CQs. CQ2Onto measures full-ontology generation over 118 CQs across five targets.

6 domain ontologies CQ2Term over 99 CQs CQ2Onto over 118 CQs 9 LLMs baseline 3 generation strategies

Explore the benchmark

Three ways in. See how models score on the leaderboard, read the formal task, dimension, and metric definitions in the catalogue, or browse the six source ontologies and their competency questions.

Cite this benchmark

CQ4OE is developed by Jiayi Li, Ziyuan Wang, Daniel Garijo, and María Poveda-Villalón at the Ontology Engineering Group, Universidad Politécnica de Madrid. If you use it in your research, please cite the archived release.

DOI: 10.5281/zenodo.20080309 Dataset: HuggingFace License: Apache-2.0
@software{cq4oe2026,
  title  = {CQ4OE: A benchmark for assessing LLM-assisted ontology
            generation from competency questions},
  author = {Li, Jiayi and Wang, Ziyuan and Garijo, Daniel and
            Poveda-Villal\'on, Mar\'ia},
  year   = {2026},
  doi    = {10.5281/zenodo.20080309},
  url    = {https://doi.org/10.5281/zenodo.20080309}
}

Acknowledgements

CQ4OE is maintained by the Ontology Engineering Group (OEG) at Universidad Politécnica de Madrid. This work was supported by the grant SOEL (Supporting Ontology Engineering with Large Language Models), PID2023-152703NA-I00, funded by MCIN/AEI/10.13039/501100011033 and by ERDF/UE.