Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
160
datasets available to search
ShareScore release 0.7.1
Dataset results
160 results for “legal”
Survey Results for: MM-LDTF: Multi-Model Legal Document Translation Framework for Vernacular Languages
Open the record for dataset details and reuse information.
Distribution. Turkey, S Russia (Daghestan and Chechnya), NE Georgia, Armenia, Azerbaijan, N Iraq, Iran, Turkmenistan, Afghanistan, and SW Pakistan. A free-roaming population introduced in 1970 in SC New Mexico, USA, increased to 2000 animals, but the population is actually maintained at 500-1000 by legalized sport hunting. in Bovidae
Distribution. Turkey, S Russia (Daghestan and Chechnya), NE Georgia, Armenia, Azerbaijan, N Iraq, Iran, Turkmenistan, Afghanistan, and SW Pakistan. A free-roaming population introduced in 1970 in SC New Mexico, USA, increased to 2000 animals, but the population is actually maintained at 500-1000 by legalized sport hunting.
Romanian Named Entity Recognition in the Legal domain (LegalNERo)
<p>LegalNERo is a manually annotated corpus for named entity recognition in the Romanian legal domain. <br> It provides gold annotations for organizations, locations, persons, time and legal resources mentioned in legal documents (legal references).Starting with version 4, the legal references were annotated using fine-grained legal document types: Law. Ordinance, Publication, Decree, Decision, Treaty, Report, Order, Regulation, Directive, EmergencyOrdinance, Norm, Convention, Code, and Other.<br> Additionally it offers GEONAMES codes for the named entities annotated as location (where a link could be established). </p> <p>The LegalNERo corpus is available in different formats: span-based, token-based and RDF. <br> The Linguistic Linked Open Data (LLOD) version is provided in RDF-Turtle format.</p> <p>CONLLUP files conform to the CoNLL-U Plus format <a href="https://universaldependencies.org/ext-format.html">https://universaldependencies.org/ext-format.html</a> .<br> Part-of-speech tagging was realized using UDPIPE. <br> Named entity annotations are placed in the column "RELATE:NE" (the 11th column) as defined in the "global.columns" metadata field.<br> Similarly GEONAMES references are in the column "RELATE:GEONAMES" (the 12th column, last).<br> Automatic processing was performed through the RELATE platform (<a href="https://relate.racai.ro">https://relate.racai.ro</a>).</p> <p>ANN files conform to BRAT format (<a href="https://brat.nlplab.org/">https://brat.nlplab.org/</a>).<br> <br> The archive contains: </p> <p>- ann_LEGAL_PER_LOC_ORG_TIME_overlap <br> Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations and time. <br> Overlapping annotations of organizations and time entities inside legal references were allowed. </p> <p>- ann_FGLEGAL_PER_LOC_ORG_TIME_overlap <br> Folder (corresponding to the above entry) in which all the files are in .ann format and contains annotations of: fine-grained legal references, persons, locations, organizations and time. <br> Overlapping annotations of organizations and time entities inside legal references were allowed. </p> <p>- ann_LEGAL_PER_LOC_ORG_TIME <br> Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. </p> <p>- ann_FGLEGAL_PER_LOC_ORG_TIME <br> Folder (corresponding to the above entry) in which all the files are in .ann format and contains annotations of: fine-grained legal references, persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. </p> <p>- ann_PER_LOC_ORG_TIME <br> Folder in which all the files are in .ann format and contains annotations of: persons, locations, organizations and time. <br> There are no overlapping annotations. </p> <p>- conllup_LEGAL_PER_LOC_ORG_TIME <br> Folder in which all the files are in .conllup format and contains annotations of: legal references, persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. <br> The annotation of these files was enhanced with GEONAMES codes (where linking was possible). </p> <p>- conllup_FGLEGAL_PER_LOC_ORG_TIME <br> Folder (corresponding to the above entry) in which all the files are in .conllup format and contains annotations of: fine-grained legal references, persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. <br> The annotation of these files was enhanced with GEONAMES codes (where linking was possible). </p> <p>- conllup_PER_LOC_ORG_TIME <br> Folder in which all the files are in .conllup format and contains annotations of: persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. <br> The annotation of these files was enhanced with GEONAMES codes (where linking was possible).</p> <p>- rdf <br> Folder containing the corpus in RDF-Turtle format.<br> All the annotations are available here in both span and token format.</p> <p>- text <br> Folder containing the raw texts.</p> <p>- splits_FGLEGAL_PER_LOC_ORG_TIME.tsv<br> This is a proposed split of the documents for training a NER system using the fine-grained entity classes. <br> The split was created randomly, while trying to ensure 15% of each entity type for validation, 15% for testing and 70% for training.<br> </p> <p><strong>NER System</strong></p> <p>A NER model generated using the LegalNERo corpus can be used online in the RELATE platform: https://relate.racai.ro/index.php?path=ner/demo</p> <p>This system was described in: Păiș, Vasile and Mitrofan, Maria and Gasan, Carol Luca and Coneschi, Vlad and Ianov, Alexandru. Named Entity Recognition in the Romanian Legal Domain. In Proceedings of the Natural Legal Language Processing Workshop 2021. Association for Computational Linguistics, Punta Cana, Dominican Republic, pp. 9--18, nov 2021</p> <p><br> <strong>LICENSING</strong></p> <p>This work is provided under the license CC BY-NC-ND 4.0 (Attribution-NonCommercial-NoDerivatives 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-nd/4.0/ <br> and the full text here: https://creativecommons.org/licenses/by-nc-nd/4.0/legalcode . </p> <p><br> <strong>CONTACT</strong></p> <p>Research Institute for Artificial Intelligence "Mihai Draganescu", Romanian Academy<br> Web: <a href="http://www.racai.ro">http://www.racai.ro</a> <br> Contact emails: vasile@racai.ro , maria@racai.ro</p>
Survey dataset for project "US-based legal interpreters and their use of academic research"
<p>Dataset of article entitled <span>US-based legal interpreters and their use of academic research: expectations, practices, and potential for collaboration between academia and the profession. Includes survey questions (PDF) and survey raw data (Excel, .xlsx file)</span></p>
Screening of legal, strategic and regulatory framework for Energy Sector - North Macedonia
<p>The dataset represents a screening of legal, strategic and regulatory framework for Energy Sector in North Macedonia. The screening contributes to the results of the research work ' <strong><em>Assessing the state and impact towards Just Transition process in the energy sector in North Macedonia – with territorial focus on the Southwest planning region'</em></strong><em>, conducted by the authors. </em></p> <p>The research is designed in the framework of the <a href="https://greenforcetwinning.net/">GreenFORCE project</a>, and will be published as a project deliverable in <em>D4.5_1st Research Study Report</em>; and <em>D4.6_2nd Research Study Report</em>.</p>
Fake legal logging in the Brazilian Amazon
<p>Abstract paper:<br> Declining deforestation rates in the Brazilian Amazon are touted as a conservation success, but illegal logging is a problem of similar scale. Recent regulatory efforts have improved detection of some forms of illegal logging, but are vulnerable to more subtle methods that mask the origin of illegal timber. We analyzed discrepancies between estimated timber volumes of official forest inventories and volumes declared in logging permits, as an indicator of potential fraud in the timber industry in the eastern Amazon. We found a strong overestimation bias of high-value timber species volumes in logging permits. Field assessments confirmed fraud for the most valuable species and complementary strategies to generate a “surplus” of licensed timber that can be used to legalize the timber coming from illegal logging. We advocate for changes to the logging control system to prevent overexploitation of Amazonian timber species and the widespread forest degradation associated with illegal logging.</p>
ASSESSING THE IMPACT OF AGROINDUSTRY AND EXTRACTIVISM ON DEFORESTATION IN THE BRAZILIAN LEGAL AMAZON
<p>Database to investigate the interrelationship between deforestation in the Legal Amazon, agroindustry and extractivism in the period from 1988 to 2022.</p>
1ra sesión - Curso "Informe Psicológico y su aplicabilidad en la Ley 348" - Módulo 1: Fundamentos y diferenciación de informe clínico, informe forense en el ámbito legal
Open the record for dataset details and reuse information.
Supplementary material for Evaluating Legal Compliance of Smart Contracts Generated by Large Language Models
<p>This repository contains the supplementary material for the paper titled "Evaluating Legal Compliance of Smart Contracts Generated by Large Language Models". It includes natural-language legal contracts, their smart contract implementations, and Petri net models of said legal contracts contracts.</p>
MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
<p>The dataset is published with:<br> <br> <em>MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer. Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Punta Cana, Dominican Republic.</em><br> <br> <strong>Documents: </strong>MultiEURLEX comprises 65k EU in 23 official EU languages. Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU. Each EUROVOC label ID is associated with a Label descriptor, e.g., [60, `agri-foodstuffs'], [6006, `plant product'], [1115, `fruit']. The descriptors are also available in 23 languages. Chalkidis et al. (2019) published a monolingual (English) version of this dataset, called EURLEX57K, comprising 57k EU laws with the originally assigned gold labels.</p> <p><strong>Languages: </strong>MultiEURLEX covers 23 languages from 7 families. EU laws are published in all official EU languages, except for Irish for resource-related reasons (Read more: https://europa.eu/european-union/about-eu/eu-languages_en). This wide coverage makes the dataset a valuable testbed for cross-lingual transfer. All languages use the Latin script, except for Bulgarian (Cyrillic script) and Greek.</p> <p><strong>Multi-granular Labeling: </strong>EUROVOC<strong> </strong>has eight levels of concepts. Each document is assigned one or more concepts (labels). If a document is assigned a concept, the ancestors and descendants of that concept are typically not assigned to the same document. The documents were originally annotated with concepts from levels 3 to 8. We created three alternative sets of labels per document, by replacing each assigned concept by its ancestor from levels 1, 2, or 3, respectively. Thus, we provide four sets of gold labels per document, one for each of the first three levels of the hierarchy, plus the original sparse label assignment.</p> <p><strong>Supported Tasks: </strong>Similarly to EURLEX (Chalkidis et al., 2019), MultiEURLEX can be used for legal topic classification, a multi-label classification task where legal documents need to be assigned concepts (in our case, from EUROVOC) reflecting their topics. Unlike EURLEX57K, however, MultiEURLEX supports labels from three different granularities (EUROVOC levels). More importantly, apart from monolingual (one-to-one) experiments, it can be used to study cross-lingual transfer scenarios, including one-to-many (systems trained in one language and used in other languages with no training data), and many-to-one or many-to-many (systems jointly trained in multiple languages and used in one or more other languages).</p> <p><strong>Data Split and Concept Drift: </strong>MultiEURLEX is chronologically split in training (55k, 1958-2010), development (5k, 2010-2012), test (5k, 2012-2016) subsets, using the English documents. The test subset contains the same 5k documents in all 23 languages. The development subset also contains the same 5k documents in 23 languages, except Croatian. Croatia is the most recent EU member (2013); older laws are gradually translated. For the official languages of the seven oldest member countries, the same 55k training documents are available; for the other languages, only a subset of the 55k training documents is available. Compared to EURLEX57K (Chalkidis et al., 2019), MultiEURLEX is not only larger (8k more documents) and multilingual; it is also more challenging, as the chronological split leads to temporal real-world concept drift across the training, development, test subsets, i.e., differences in label distribution and phrasing, representing a realistic temporal generalization problem (Huang and Paul, 2019; Lazaridou et al., 2021). Recently, Søgaard et al. (2021) showed this setup is more realistic, as it does not overestimate real performance, contrary to random splits (Gorman and Bedrick, 2019).</p>
Greek_Legal_Code
<p>Greek_Legal_Code contains 47k classified legal resources from Greek Legislation. Its origin is “Permanent Greek Legislation Code - Raptarchis”, a collection of Greek legislative documents classified into multi-level (from broader to more specialized) categories. Format: jsonl</p>
LAW, LOGIC AND STRUCTURATION IN LEGAL SYSTEMATICS
<p>Presentation at the Department of Theory and History of State and Law of St. Petersburg State University conference on »Systematization in Law: ʻMagic Glassʼ of the Codifier (to the 250th anniversary of the birth of Mikhail Mikhailovich Speransky)« organised on October 14, 2022 [19:12]</p> <p>1. Theoretical Background 2. Foundations of Structuration and Systematisation Challenged 3. Is there a Structure? 4. Structuring as a Meta-construct</p>
VULNER legal and policy documents
<p>Legal and policy documents, which were collected as part of the VULNER project.</p>
Ditte Laursen: Developing a legal agreement for research-access to a web archive
<p>Ditte Laursen: Developing a legal agreement for research-access to a web archive</p> <p>WARcnet Luxembourg meeting Thursday 5 November 2020</p>
Dataset for Thesis "Automated Consistency of Legal and Software Architecture System Specifications for Data Protection Analysis"
<p>Dataset for Thesis "Automated Consistency of Legal and Software Architecture System Specifications for Data Protection Analysis"</p>
Expanding the UTHealth Medical Legal Partnership to Improve Mental Health for Low-Income Individuals
ClinicalTrials.gov study NCT03805126. IPD Sharing: NO. Countries: 1. Publications: 1.
Juvenile Justice Translational Research on Interventions for Adolescents in the Legal System
ClinicalTrials.gov study NCT02672150. IPD Sharing: YES. Countries: 1. Publications: 8.
Evaluation of of a Prefixed 50% N2O- 50%O2 Mixture in Legal Abortion Under Local Analgesia
ClinicalTrials.gov study NCT00769912. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Recovery Legal Care Clinical Trial
ClinicalTrials.gov study NCT06618794. IPD Sharing: YES. Countries: 1. Publications: 0.
Pharmacologic Treatment in Legal Offenders With Schizophrenia, a Prospective Observational Mirror Image Study.
ClinicalTrials.gov study NCT05939765. IPD Sharing: NO. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.