Dataset Overview
The NL2SHACL dataset supports evaluation of natural language to SHACL translation across diverse domains. It comprises six subsets covering both domain-specific and general knowledge settings.
Each record follows a unified structure consisting of a natural language description, a reference SHACL shapes graph, and an ontology snippet capturing the semantics of the terms referenced in the shapes graph.

Statistics
| Subset | # Records | # Node Shapes | # Prop. Shapes | Avg. NL Length | Avg. Shapes/Record |
|---|---|---|---|---|---|
| CHEMROF | 74 | 74 | 441 | 154.5 | 6.96 |
| DCAT | 20 | 35 | 103 | 116.2 | 6.90 |
| ePO | 50 | 50 | 139 | 116.5 | 3.78 |
| Invoice | 78 | 78 | 113 | 75.1 | 2.45 |
| SNIK | 8 | 11 | 49 | 52.0 | 7.50 |
| DBpedia | 10 | 10 | 11 | 9.5 | 2.10 |
| Overall | 240 | 258 | 856 | 108.1 | 4.64 |
Filtering Summary
| Subset | # Raw Records | # Filtered Records |
|---|---|---|
| CHEMROF | 128 | 17 |
| DCAT | 21 | 1 |
| ePO | 378 | 235 |
| Invoice | 84 | 6 |
| SNIK | 27 | 0 |
| DBpedia | -- | -- |
Records are filtered during preprocessing due to syntactic invalidity, unresolved ontology terms, or SPARQL-based constraints. See the Framework page for details on the filtering process.
Data Sources
Four subsets (CHEMROF, DCAT, ePO, SNIK) are selected from the Shapes of You index. The Invoice subset is derived from prior work on EDIFACT-based invoice validation. The DBpedia subset is curated by the authors using the DBpedia ontology.
The subsets cover a diverse range of constraint patterns, from structural and datatype validation to semantic and logical constraints.