Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

160

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

160 results for “legal”

Learn how ShareScore rates datasets ↗
zenodo40/100

FIGURE 4 in Reproductive biology of seven fish species of commercial interest at the Ramsar site in the Baixada Maranhense, Legal Amazon, Brazil

FIGURE 4 | Gonadosomatic index indicating the spawning season for (A) Cichla monoculus; (B) Hassar affinis; (C) Hoplias malabaricus; (D) Plagioscion squamosissimus; (E) Prochilodus lacustris; (F) Pygocentrus nattereri; and (G) Schizodon dissimilis, caught in Baixada Maranhense Protection Area, between January 2012 and December 2016.

opencc-by-4.0Jul 2021View details →
zenodo40/100

Spanish Legal Domain Word & Sub-Word Embeddings

<p><strong>Spanish Legal Word and Sub-word Embeddings in FastText</strong></p> <p>These embeddings have been generated from the largest corpus (9GB)&nbsp;ever made from Spanish Legal resources till the date.</p> <p>More legal domain resources:&nbsp;https://github.com/PlanTL-GOB-ES/lm-legal-es</p> <p><strong>Citation</strong></p> <pre><code>@misc{gutierrezfandino2021legal, title={Spanish Legalese Language Model and Corpora}, author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Aitor Gonzalez-Agirre and Marta Villegas}, year={2021}, eprint={2110.12201}, archivePrefix={arXiv}, primaryClass={cs.CL} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Spanish Legal Domain Corpora

<p><strong>Spanish Legal Domain Corpora</strong></p> <p>A collection of corpora of Spanish legal domain.</p> <p>More legal domain resources:&nbsp;https://github.com/PlanTL-GOB-ES/lm-legal-es</p> <p><strong>Citation</strong></p> <pre><code>@misc{gutierrezfandino2021legal, title={Spanish Legalese Language Model and Corpora}, author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Aitor Gonzalez-Agirre and Marta Villegas}, year={2021}, eprint={2110.12201}, archivePrefix={arXiv}, primaryClass={cs.CL} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Analysis of the concept of informal economy through 102 definitions: legality or necessity

<p><strong>Abstract</strong></p> <p>The processes of informal economy are well established, but the same cannot be said of their conceptual treatment in the academic literature. They constitute complex phenomena that cut across sectors and disciplines and give rise to other elements that simultaneously reject and encourage them. For many formal stakeholders in the economy, they are an enemy to be beaten; for the authorities, informal activity is seen as a loss of revenue for the state coffers; for the Sustainable Development Goals, by implicitly recognizing them in goal 8, they constitute a paradigm shift.&nbsp; Meanwhile, the reality for those involved in the informal economy is that it is a way of life and not a mere choice, one that leads to the most social of all economies: that of necessity. There is no consensus among academics on informality and its ramifications, hence the need to analyze the processes of informal economy from its theoretical construction with the purpose of discovering its range and depth, as well as its interrelationships and theoretical implications. To achieve this, 102 definitions of informal economy were analyzed by identifying and deconstructing their dimensions and performing a frequency count of their citation in Google Scholar. This analysis demonstrated the lack of cultural elements in the definitions, which are the true underlying cause of the phenomenon, and the over-prominence afforded to legal dimensions.</p> <p><strong>Plain language summary</strong></p> <p>People have to find many different ways to earn a living in order to meet the needs of their families.&nbsp; Some of these modes of employment, the so-called informal economy, clash directly with the wishes of governments and organizations who do not recognize that decent paid work may take many forms or do not understand that there are often no alternatives. This confused picture of contradictory opinions requires some clarity, which was the aim of this study in its analysis of the definition of informal economy by 102 experts. The analysis sought to reveal the meaning of this concept that is fundamental to life in society, especially in countries suffering from environmental challenges, political crises, or the scourge of corruption. The findings demonstrate the lack of cultural considerations in the academic study of this complex phenomenon, despite their inherent importance. If progress is to be made in improving working conditions, greater understanding is needed of how the most basic needs affect the way society functions.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

<p>This benchmark dataset is published with the article:&nbsp;</p> <p><em>Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras. 2021.&nbsp;LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. ArXiv.</em></p> <p><strong>Short Description</strong></p> <p>Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela,2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., &nbsp;2019), we introduce LexGLUE, a benchmark dataset to evaluate the performance of NLP methods in legal tasks. LexGLUE is based on seven existing legal NLP datasets:</p> <ul> <li>ECtHR Task A &nbsp;(Chalkidis et al., 2019)</li> <li>ECtHR Task B &nbsp;(Chalkidis et al., 2021a)</li> <li>SCOTUS (Spaeth et al., 2020)</li> <li>EUR-LEX (Chalkidis et al., 2021b)</li> <li>LEDGAR (Tuggener et al. (2020)</li> <li>UNFAIR-ToS (Lippi et al., 2019)</li> <li>CaseHOLD (Zheng et al., 2021)</li> </ul>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Systematic Mapping Study in Legal NLP (Raw Data)

<p>Systematic Mapping Study in Legal NLP (Raw Data). It contains all papers processed, accepted, information extraction and all the steps taken, and which researchers reviewed what.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

JusBrasilRec: A large-scale dataset of user sessions for recommendations on the legal domain

<p><strong>JusBrasilRec: A large-scale dataset of user sessions for recommendations on the legal domain</strong></p> <p>The proliferation of legal documents in various formats and their dispersion across multiple courts present a significant challenge for users seeking precise matches to their information requirements. Despite notable advancements in legal information retrieval systems, research into legal recommender systems remains limited. A plausible factor contributing to this scarcity could be the absence of extensive publicly accessible datasets or benchmarks.</p> <p>Jusbrasil (<a href="https://www.jusbrasil.com.br">https://www.jusbrasil.com.br</a>) is known as the largest legal search portal in Brazil. It provides an online environment where users can find the legal documents that best match their information needs. With millions of user interactions to billions of documents containing different artifacts related to law in Brazil, Jusbrasil appears as a large-scale test bed for advancing research on the still scarce area of legal recommender systems.&nbsp;</p> <p>Therefore, we&nbsp;collected and made available the <strong>JusBrasilRec</strong>, a dataset containing user sessions from Jusbrasil for recommendations on the legal domain. Additionally, we also computed and made available a TF-IDF matrix from the textual content of the documents in Jusbrasil. The following files are available for download from JusBrasilRec:</p> <ul> <li><strong>jusbrasilrec_dataset.zip:</strong> a compacted file containing the user sessions;</li> <li><strong>jusbrasilrec_tfidf_matrix.zip:</strong> a compacted file containing the TF-IDF matrix;</li> <li><strong>readme.txt:</strong> a text file explaining the content and format of the previous files.</li> </ul> <p><strong>How to cite the dataset:</strong> Marcos Aur&eacute;lio Domingues, Edleno Silva de Moura, Leandro Balby Marinho and Altigran da Silva. A Large Scale Benchmark for Session-based Recommendations on the Legal Domain. Artificial Intelligence and Law. 2023.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

The Legal Spectrum

<p>The Legal Spectrum shows legal information on the range from closed to open. While contracts (or other voluntarily entered agreements) are often closed or shared, legislation and case law are - or at least should be - available to the general public without restrictions. </p> <p>The Legal Spectrum is based on the Data Spectrum, originally published by the Open Data Institute (ODI).</p>

opencc-by-sa-4.0Oct 2016View details →
dryad36/100

Are cryptocurrencies currencies? Bitcoin as legal tender in El Salvador

<p>A currency's essential feature is to be a medium of exchange. We leverage a quasi-natural experiment––El Salvador as the first country to make Bitcoin legal tender––to study a cryptocurrency's potential to be used in daily transactions. The government also launched and provided incentives to download and use a digital wallet named Chivo, which shares features with Central Bank Digital Currencies (CBDCs) and allows users to trade bitcoins and dollars. Were Chivo Wallet and Bitcoin actually adopted after this "big push"? Conducting a representative face-to-face survey and relying on blockchain data to obtain all Chivo transactions, we document how usage of digital payments and Bitcoin is low, concentrated, and has been decreasing over time. We find that privacy concerns are key barriers to adoption, which speaks to a policy debate on crypto and CBDCs that has had anonymity at its core. We also estimate the technology's adoption cost and its network externalities.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Arabic Handwritten Legal Amount (AHLA) Dataset

<p>The AHLA dataset is collected by distributing an advanced designed report with Arabic native speakers. Our dataset contains two kinds of Arabic handwritten :&nbsp;<br>(1) &nbsp; &nbsp;Arabic word-level images that express legal amounts of bank cheques, including the colloquial words used in writing Arabic numbers.&nbsp;<br>(2) &nbsp; &nbsp;Arabic legal amount sentence images.</p> <p>The primary objective of compiling this comprehensive dataset is to furnish a diverse range of Arabic language samples. These samples are intended for training and testing systems capable of autonomously recognizing and comprehending handwritten legal amounts on financial documents. Subsequently, the aim is to convert these semantic expressions into their respective numeric currency totals, facilitating digital processing and banking operations.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Additional dataset to Trading deforestation—why the legality of forest-risk commodities is insufficient

<p>This dataset refers to the summary of unprotected native vegetation and carbon stocks per municipality in Brazil. Unprotected means native vegetation areas and spatially-linked carbon stocks present at farms with surplus of legal reserves according to the Brazilian Forest Code.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Dataset and additional files/softwares required for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents"

<p>This dump contains all files and softwares required for running the codes for the paper&nbsp;&quot;LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents&quot;. Specifically, these codes are available at&nbsp;https://github.com/Law-AI/LeSICiN.</p> <p>LeSICiN is a deep neural network for the task of Legal Statute Identification which also uses graphical properties of the document-statute citation network for training and predictions.</p> <p>We have three datasets --- train, dev and test. These are all .jsonl files with each instance dict per line; each instance dict contains the unique id, list of sentences and cited labels of the particular instance. Also, there is a fourth file --- secs.jsonl, which stores the text of all the statutes in similar format.</p> <p>schemas.json list out the metapath schemas for fact and section type nodes, while type_map.json maps the id of each node to its type (Act/Chapter/Topic/Section/Fact).&nbsp;</p> <p>label_tree.json and citation_network.json list out the edges for the two parts of the network in the format of a 3-tuple (&#39;source id&#39;, &#39;relationship type&#39;, &#39;target id&#39;)</p> <p>&quot;ils2v.bin&quot; is the pretrained sent2vec vectorizer that can generate a 200-dim vector for each sentence</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Fairlex: A multilingual benchmark for evaluating fairness in legal text processing

<p>We present a benchmark suite of four datasets for evaluating the fairness of pre-trained legal language models and the techniques used to fine-tune them for downstream tasks.&nbsp;Our benchmarks cover four jurisdictions&nbsp;(European Council,&nbsp;USA,&nbsp;Swiss,&nbsp;and Chinese),&nbsp;five languages&nbsp;(English,&nbsp;German,&nbsp;French,&nbsp;Italian, and Chinese), and fairness across five attributes&nbsp;(gender,&nbsp;age,&nbsp;nationality/region,&nbsp;language,&nbsp;and legal area).&nbsp;In our experiments,&nbsp;we evaluate pre-trained language models using several group-robust fine-tuning techniques and show that performance group disparities are vibrant in many cases,&nbsp;while none of these techniques guarantee fairness,&nbsp;nor consistently mitigate group disparities.&nbsp;Furthermore,&nbsp;we provide a quantitative and qualitative analysis of our results,&nbsp;highlighting open challenges in the development of robustness methods in legal NLP.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Dominant Rural Technological Trajectories (TTs) dataset at municipality level of the Brazilian Legal Amazon (BLA)

<p>This dataset contains the dominant technological trajectories (TTs) of the Brazilian Legal Amazon (BLA) municipalities for the years of 1995, 2006 and 2017. The dominant trajectory is the one, among the six identified by Costa (2021), that is economically most important in the municipality. The relative share of the Gross Value of Rural Production of the trajectory in the total Gross Value of Rural Production in the municipality was taken as a proxy of economic importance. From the tabulation of the datasets in Costa (2022), the dominant technological trajectory was identified, it means, considering the methodology used, which of the six TTs was responsible for over 50% of the municipal Gross Value of Rural Production. The dominant TTs were calculated using the official municipal grid for the year the agricultural census was carried out.</p> <p>The dataset is organized as a .csv table, for each year (1995, 2006 and 2017), with a geographic key for each municipality (6-digit municipality code, 2-digit state code), that can be easily linked with municipality available shapefiles and other datasets.&nbsp;</p> <p>References:</p> <p>Costa, F. A. Structural diversity and change in rural Amazonia: a comparative assessment of the technological trajectories based on agricultural censuses (1995, 2006 and 2017). <strong>Nova econ</strong>. 31 (02), May-Aug 2021, doi:10.1590/0103-6351/6373.</p> <p>Costa, F. A, (2022). Database of Rural Technological Trajectories of the Legal Amazon delimited by the Method of Differentiation and Structural Signification of Rural Production (M-DASTRU). <strong>Zenodo.</strong> DOI: 10.5281/zenodo.7035753</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation

<p>This repository contains the following 3&nbsp;datasets for legal document summarization :</p> <p>- IN-Abs : Indian Supreme Court case documents &amp; their `abstractive&#39; summaries, obtained from http://www.liiofindia.org/in/cases/cen/INSC/<br> - IN-Ext : Indian Supreme Court case documents &amp; their `extractive&#39; summaries, written by two law experts (A1, A2).<br> - UK-Abs : United Kingdom (U.K.) Supreme Court case documents &amp; their `abstractive&#39; summaries, obtained from https://www.supremecourt.uk/decided-cases/</p> <p>Please refer to the paper and the README file for more details.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Supplementary Material for A Metrics Suite for Quantifying Legal Compliance of Smart Contracts

<p>This repository contains the supplementary material for the paper titled "A Metrics Suite for Quantifying Legal Compliance of Smart Contracts". It includes natural-language legal contracts, their smart contract implementations, Petri net models of said contracts, and their reachability graphs. The Petri net models are presented as graphics, as well as .cpn files that can be opened with either <a href="https://cpntools.org/" target="_blank" rel="noopener">CPN Tools</a> or <a href="https://cpnide.org/" target="_blank" rel="noopener">CPN IDE</a>.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Fig. 6 in Identification of Muscidae (Diptera) of medico-legal importance by means of wing measurements

Fig. 6 Discrimination of four species of the Muscina based on canonical variate analysis

opencc-by-4.0Mar 2017View details →
zenodo36/100

Fig. 5 in Identification of Muscidae (Diptera) of medico-legal importance by means of wing measurements

Fig. 5 Discrimination of eight species of the Hydrotaea based on canonical variate analysis

opencc-by-4.0Mar 2017View details →
zenodo36/100

Fig. 4 in Identification of Muscidae (Diptera) of medico-legal importance by means of wing measurements

Fig. 4 Discrimination of Muscidae genera based on canonical variate analysis

opencc-by-4.0Mar 2017View details →
zenodo36/100

Sistema para la gestion de expedinetes en el area Legal

<p><strong><span>RESUMEN</span></strong></p> <p>El caso de estudio se centr&oacute; en el desarrollo e implementaci&oacute;n de un sistema web dise&ntilde;ado espec&iacute;ficamente para la gesti&oacute;n de expedientes en una consultor&iacute;a legal especializada en derecho de familia. El objetivo principal del sistema era optimizar la organizaci&oacute;n, acceso y seguimiento de la informaci&oacute;n relacionada con los casos de los clientes.</p> <p>El sistema web permite a los abogados y al personal administrativo de la consultor&iacute;a almacenar y gestionar los expedientes de manera digital, lo que mejoraba la eficiencia y reduc&iacute;a la dependencia de documentos f&iacute;sicos. Adem&aacute;s, ofrece funcionalidades como la programaci&oacute;n de citas, la gesti&oacute;n de tareas y recordatorios, y la generaci&oacute;n de informes para facilitar el seguimiento y la evaluaci&oacute;n del progreso de cada caso.</p> <p>El desarrollo del sistema web implic&oacute; un proceso de an&aacute;lisis detallado de los requisitos espec&iacute;ficos de la consultor&iacute;a, seguido de la fase de dise&ntilde;o e implementaci&oacute;n por parte de un equipo de desarrolladores especializados en tecnolog&iacute;as web. Se realizaron pruebas exhaustivas para garantizar la fiabilidad y seguridad del sistema antes de su implementaci&oacute;n final.</p> <p>Tras la implementaci&oacute;n del sistema web, se observaran mejoras significativas en la productividad y eficiencia del equipo legal. La capacidad de acceder r&aacute;pidamente a la informaci&oacute;n relevante y de colaborar de manera m&aacute;s efectiva contribuy&oacute; a una prestaci&oacute;n de servicios m&aacute;s &aacute;gil y de mayor calidad para los clientes de la consultor&iacute;a.</p> <p>En resumen, el caso de estudio destac&oacute; c&oacute;mo la implementaci&oacute;n de un sistema web personalizado puede transformar y mejorar la gesti&oacute;n de expedientes en una consultor&iacute;a legal, optimizando los procesos internos y mejorando la experiencia del cliente.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record