This is training data for machine learning models that can be used with Cavil,
the openSUSE legal review and SBOM system.
Cavil uses a pattern matching system to identify potential legal text in source code. This process is based around identifying
hot zones of legal keywords (snippets) and produces around 80% false positives. Historically these false positives had to be
sorted out by humans. A few years ago we've started using machine learning to automate much of… See the full description on the dataset page:
https://huggingface.co/datasets/openSUSE/cavil-legal-text.