Skip to content

Dataset Overview

The NL2SHACL dataset supports evaluation of natural language to SHACL translation across diverse domains. It comprises six subsets covering both domain-specific and general knowledge settings.

Each record follows a unified structure consisting of a natural language description, a reference SHACL shapes graph, and an ontology snippet capturing the semantics of the terms referenced in the shapes graph.

Example Record


Statistics

Subset # Records # Node Shapes # Prop. Shapes Avg. NL Length Avg. Shapes/Record
CHEMROF 74 74 441 154.5 6.96
DCAT 20 35 103 116.2 6.90
ePO 50 50 139 116.5 3.78
Invoice 78 78 113 75.1 2.45
SNIK 8 11 49 52.0 7.50
DBpedia 10 10 11 9.5 2.10
Overall 240 258 856 108.1 4.64

Filtering Summary

Subset # Raw Records # Filtered Records
CHEMROF 128 17
DCAT 21 1
ePO 378 235
Invoice 84 6
SNIK 27 0
DBpedia -- --

Records are filtered during preprocessing due to syntactic invalidity, unresolved ontology terms, or SPARQL-based constraints. See the Framework page for details on the filtering process.


Data Sources

Four subsets (CHEMROF, DCAT, ePO, SNIK) are selected from the Shapes of You index. The Invoice subset is derived from prior work on EDIFACT-based invoice validation. The DBpedia subset is curated by the authors using the DBpedia ontology.

The subsets cover a diverse range of constraint patterns, from structural and datatype validation to semantic and logical constraints.