← Home

CQ4OE Benchmark Leaderboard

CQ4OE leaderboard for LLM-based ontology generation from competency questions. Runs are grouped by evaluation configuration so only identical settings are ranked against each other. The Status column tells you who has checked a run. Reproduced: we ran the model ourselves and got the same kind of output. Verified: we also went through the alignment by hand and corrected any mistakes before scoring. Unchecked runs still appear, but only runs with both badges get a rank.