Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
160
datasets available to search
ShareScore release 0.7.1
Dataset results
160 results for “legal”
FIGURE 4 in Reproductive biology of seven fish species of commercial interest at the Ramsar site in the Baixada Maranhense, Legal Amazon, Brazil
FIGURE 4 | Gonadosomatic index indicating the spawning season for (A) Cichla monoculus; (B) Hassar affinis; (C) Hoplias malabaricus; (D) Plagioscion squamosissimus; (E) Prochilodus lacustris; (F) Pygocentrus nattereri; and (G) Schizodon dissimilis, caught in Baixada Maranhense Protection Area, between January 2012 and December 2016.
Spanish Legal Domain Word & Sub-Word Embeddings
<p><strong>Spanish Legal Word and Sub-word Embeddings in FastText</strong></p> <p>These embeddings have been generated from the largest corpus (9GB) ever made from Spanish Legal resources till the date.</p> <p>More legal domain resources: https://github.com/PlanTL-GOB-ES/lm-legal-es</p> <p><strong>Citation</strong></p> <pre><code>@misc{gutierrezfandino2021legal, title={Spanish Legalese Language Model and Corpora}, author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Aitor Gonzalez-Agirre and Marta Villegas}, year={2021}, eprint={2110.12201}, archivePrefix={arXiv}, primaryClass={cs.CL} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Spanish Legal Domain Corpora
<p><strong>Spanish Legal Domain Corpora</strong></p> <p>A collection of corpora of Spanish legal domain.</p> <p>More legal domain resources: https://github.com/PlanTL-GOB-ES/lm-legal-es</p> <p><strong>Citation</strong></p> <pre><code>@misc{gutierrezfandino2021legal, title={Spanish Legalese Language Model and Corpora}, author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Aitor Gonzalez-Agirre and Marta Villegas}, year={2021}, eprint={2110.12201}, archivePrefix={arXiv}, primaryClass={cs.CL} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Analysis of the concept of informal economy through 102 definitions: legality or necessity
<p><strong>Abstract</strong></p> <p>The processes of informal economy are well established, but the same cannot be said of their conceptual treatment in the academic literature. They constitute complex phenomena that cut across sectors and disciplines and give rise to other elements that simultaneously reject and encourage them. For many formal stakeholders in the economy, they are an enemy to be beaten; for the authorities, informal activity is seen as a loss of revenue for the state coffers; for the Sustainable Development Goals, by implicitly recognizing them in goal 8, they constitute a paradigm shift. Meanwhile, the reality for those involved in the informal economy is that it is a way of life and not a mere choice, one that leads to the most social of all economies: that of necessity. There is no consensus among academics on informality and its ramifications, hence the need to analyze the processes of informal economy from its theoretical construction with the purpose of discovering its range and depth, as well as its interrelationships and theoretical implications. To achieve this, 102 definitions of informal economy were analyzed by identifying and deconstructing their dimensions and performing a frequency count of their citation in Google Scholar. This analysis demonstrated the lack of cultural elements in the definitions, which are the true underlying cause of the phenomenon, and the over-prominence afforded to legal dimensions.</p> <p><strong>Plain language summary</strong></p> <p>People have to find many different ways to earn a living in order to meet the needs of their families. Some of these modes of employment, the so-called informal economy, clash directly with the wishes of governments and organizations who do not recognize that decent paid work may take many forms or do not understand that there are often no alternatives. This confused picture of contradictory opinions requires some clarity, which was the aim of this study in its analysis of the definition of informal economy by 102 experts. The analysis sought to reveal the meaning of this concept that is fundamental to life in society, especially in countries suffering from environmental challenges, political crises, or the scourge of corruption. The findings demonstrate the lack of cultural considerations in the academic study of this complex phenomenon, despite their inherent importance. If progress is to be made in improving working conditions, greater understanding is needed of how the most basic needs affect the way society functions.</p>
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
<p>This benchmark dataset is published with the article: </p> <p><em>Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras. 2021. LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. ArXiv.</em></p> <p><strong>Short Description</strong></p> <p>Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela,2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., 2019), we introduce LexGLUE, a benchmark dataset to evaluate the performance of NLP methods in legal tasks. LexGLUE is based on seven existing legal NLP datasets:</p> <ul> <li>ECtHR Task A (Chalkidis et al., 2019)</li> <li>ECtHR Task B (Chalkidis et al., 2021a)</li> <li>SCOTUS (Spaeth et al., 2020)</li> <li>EUR-LEX (Chalkidis et al., 2021b)</li> <li>LEDGAR (Tuggener et al. (2020)</li> <li>UNFAIR-ToS (Lippi et al., 2019)</li> <li>CaseHOLD (Zheng et al., 2021)</li> </ul>
Systematic Mapping Study in Legal NLP (Raw Data)
<p>Systematic Mapping Study in Legal NLP (Raw Data). It contains all papers processed, accepted, information extraction and all the steps taken, and which researchers reviewed what.</p>
JusBrasilRec: A large-scale dataset of user sessions for recommendations on the legal domain
<p><strong>JusBrasilRec: A large-scale dataset of user sessions for recommendations on the legal domain</strong></p> <p>The proliferation of legal documents in various formats and their dispersion across multiple courts present a significant challenge for users seeking precise matches to their information requirements. Despite notable advancements in legal information retrieval systems, research into legal recommender systems remains limited. A plausible factor contributing to this scarcity could be the absence of extensive publicly accessible datasets or benchmarks.</p> <p>Jusbrasil (<a href="https://www.jusbrasil.com.br">https://www.jusbrasil.com.br</a>) is known as the largest legal search portal in Brazil. It provides an online environment where users can find the legal documents that best match their information needs. With millions of user interactions to billions of documents containing different artifacts related to law in Brazil, Jusbrasil appears as a large-scale test bed for advancing research on the still scarce area of legal recommender systems. </p> <p>Therefore, we collected and made available the <strong>JusBrasilRec</strong>, a dataset containing user sessions from Jusbrasil for recommendations on the legal domain. Additionally, we also computed and made available a TF-IDF matrix from the textual content of the documents in Jusbrasil. The following files are available for download from JusBrasilRec:</p> <ul> <li><strong>jusbrasilrec_dataset.zip:</strong> a compacted file containing the user sessions;</li> <li><strong>jusbrasilrec_tfidf_matrix.zip:</strong> a compacted file containing the TF-IDF matrix;</li> <li><strong>readme.txt:</strong> a text file explaining the content and format of the previous files.</li> </ul> <p><strong>How to cite the dataset:</strong> Marcos Aurélio Domingues, Edleno Silva de Moura, Leandro Balby Marinho and Altigran da Silva. A Large Scale Benchmark for Session-based Recommendations on the Legal Domain. Artificial Intelligence and Law. 2023.</p>
The Legal Spectrum
<p>The Legal Spectrum shows legal information on the range from closed to open. While contracts (or other voluntarily entered agreements) are often closed or shared, legislation and case law are - or at least should be - available to the general public without restrictions. </p> <p>The Legal Spectrum is based on the Data Spectrum, originally published by the Open Data Institute (ODI).</p>
Are cryptocurrencies currencies? Bitcoin as legal tender in El Salvador
<p>A currency's essential feature is to be a medium of exchange. We leverage a quasi-natural experiment––El Salvador as the first country to make Bitcoin legal tender––to study a cryptocurrency's potential to be used in daily transactions. The government also launched and provided incentives to download and use a digital wallet named Chivo, which shares features with Central Bank Digital Currencies (CBDCs) and allows users to trade bitcoins and dollars. Were Chivo Wallet and Bitcoin actually adopted after this "big push"? Conducting a representative face-to-face survey and relying on blockchain data to obtain all Chivo transactions, we document how usage of digital payments and Bitcoin is low, concentrated, and has been decreasing over time. We find that privacy concerns are key barriers to adoption, which speaks to a policy debate on crypto and CBDCs that has had anonymity at its core. We also estimate the technology's adoption cost and its network externalities.</p>
Arabic Handwritten Legal Amount (AHLA) Dataset
<p>The AHLA dataset is collected by distributing an advanced designed report with Arabic native speakers. Our dataset contains two kinds of Arabic handwritten : <br>(1) Arabic word-level images that express legal amounts of bank cheques, including the colloquial words used in writing Arabic numbers. <br>(2) Arabic legal amount sentence images.</p> <p>The primary objective of compiling this comprehensive dataset is to furnish a diverse range of Arabic language samples. These samples are intended for training and testing systems capable of autonomously recognizing and comprehending handwritten legal amounts on financial documents. Subsequently, the aim is to convert these semantic expressions into their respective numeric currency totals, facilitating digital processing and banking operations.</p>
Additional dataset to Trading deforestation—why the legality of forest-risk commodities is insufficient
<p>This dataset refers to the summary of unprotected native vegetation and carbon stocks per municipality in Brazil. Unprotected means native vegetation areas and spatially-linked carbon stocks present at farms with surplus of legal reserves according to the Brazilian Forest Code.</p>
Dataset and additional files/softwares required for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents"
<p>This dump contains all files and softwares required for running the codes for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents". Specifically, these codes are available at https://github.com/Law-AI/LeSICiN.</p> <p>LeSICiN is a deep neural network for the task of Legal Statute Identification which also uses graphical properties of the document-statute citation network for training and predictions.</p> <p>We have three datasets --- train, dev and test. These are all .jsonl files with each instance dict per line; each instance dict contains the unique id, list of sentences and cited labels of the particular instance. Also, there is a fourth file --- secs.jsonl, which stores the text of all the statutes in similar format.</p> <p>schemas.json list out the metapath schemas for fact and section type nodes, while type_map.json maps the id of each node to its type (Act/Chapter/Topic/Section/Fact). </p> <p>label_tree.json and citation_network.json list out the edges for the two parts of the network in the format of a 3-tuple ('source id', 'relationship type', 'target id')</p> <p>"ils2v.bin" is the pretrained sent2vec vectorizer that can generate a 200-dim vector for each sentence</p>
Fairlex: A multilingual benchmark for evaluating fairness in legal text processing
<p>We present a benchmark suite of four datasets for evaluating the fairness of pre-trained legal language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA, Swiss, and Chinese), five languages (English, German, French, Italian, and Chinese), and fairness across five attributes (gender, age, nationality/region, language, and legal area). In our experiments, we evaluate pre-trained language models using several group-robust fine-tuning techniques and show that performance group disparities are vibrant in many cases, while none of these techniques guarantee fairness, nor consistently mitigate group disparities. Furthermore, we provide a quantitative and qualitative analysis of our results, highlighting open challenges in the development of robustness methods in legal NLP.</p>
Dominant Rural Technological Trajectories (TTs) dataset at municipality level of the Brazilian Legal Amazon (BLA)
<p>This dataset contains the dominant technological trajectories (TTs) of the Brazilian Legal Amazon (BLA) municipalities for the years of 1995, 2006 and 2017. The dominant trajectory is the one, among the six identified by Costa (2021), that is economically most important in the municipality. The relative share of the Gross Value of Rural Production of the trajectory in the total Gross Value of Rural Production in the municipality was taken as a proxy of economic importance. From the tabulation of the datasets in Costa (2022), the dominant technological trajectory was identified, it means, considering the methodology used, which of the six TTs was responsible for over 50% of the municipal Gross Value of Rural Production. The dominant TTs were calculated using the official municipal grid for the year the agricultural census was carried out.</p> <p>The dataset is organized as a .csv table, for each year (1995, 2006 and 2017), with a geographic key for each municipality (6-digit municipality code, 2-digit state code), that can be easily linked with municipality available shapefiles and other datasets. </p> <p>References:</p> <p>Costa, F. A. Structural diversity and change in rural Amazonia: a comparative assessment of the technological trajectories based on agricultural censuses (1995, 2006 and 2017). <strong>Nova econ</strong>. 31 (02), May-Aug 2021, doi:10.1590/0103-6351/6373.</p> <p>Costa, F. A, (2022). Database of Rural Technological Trajectories of the Legal Amazon delimited by the Method of Differentiation and Structural Signification of Rural Production (M-DASTRU). <strong>Zenodo.</strong> DOI: 10.5281/zenodo.7035753</p>
Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation
<p>This repository contains the following 3 datasets for legal document summarization :</p> <p>- IN-Abs : Indian Supreme Court case documents & their `abstractive' summaries, obtained from http://www.liiofindia.org/in/cases/cen/INSC/<br> - IN-Ext : Indian Supreme Court case documents & their `extractive' summaries, written by two law experts (A1, A2).<br> - UK-Abs : United Kingdom (U.K.) Supreme Court case documents & their `abstractive' summaries, obtained from https://www.supremecourt.uk/decided-cases/</p> <p>Please refer to the paper and the README file for more details.</p>
Supplementary Material for A Metrics Suite for Quantifying Legal Compliance of Smart Contracts
<p>This repository contains the supplementary material for the paper titled "A Metrics Suite for Quantifying Legal Compliance of Smart Contracts". It includes natural-language legal contracts, their smart contract implementations, Petri net models of said contracts, and their reachability graphs. The Petri net models are presented as graphics, as well as .cpn files that can be opened with either <a href="https://cpntools.org/" target="_blank" rel="noopener">CPN Tools</a> or <a href="https://cpnide.org/" target="_blank" rel="noopener">CPN IDE</a>.</p>
Fig. 6 in Identification of Muscidae (Diptera) of medico-legal importance by means of wing measurements
Fig. 6 Discrimination of four species of the Muscina based on canonical variate analysis
Fig. 5 in Identification of Muscidae (Diptera) of medico-legal importance by means of wing measurements
Fig. 5 Discrimination of eight species of the Hydrotaea based on canonical variate analysis
Fig. 4 in Identification of Muscidae (Diptera) of medico-legal importance by means of wing measurements
Fig. 4 Discrimination of Muscidae genera based on canonical variate analysis
Sistema para la gestion de expedinetes en el area Legal
<p><strong><span>RESUMEN</span></strong></p> <p>El caso de estudio se centró en el desarrollo e implementación de un sistema web diseñado específicamente para la gestión de expedientes en una consultoría legal especializada en derecho de familia. El objetivo principal del sistema era optimizar la organización, acceso y seguimiento de la información relacionada con los casos de los clientes.</p> <p>El sistema web permite a los abogados y al personal administrativo de la consultoría almacenar y gestionar los expedientes de manera digital, lo que mejoraba la eficiencia y reducía la dependencia de documentos físicos. Además, ofrece funcionalidades como la programación de citas, la gestión de tareas y recordatorios, y la generación de informes para facilitar el seguimiento y la evaluación del progreso de cada caso.</p> <p>El desarrollo del sistema web implicó un proceso de análisis detallado de los requisitos específicos de la consultoría, seguido de la fase de diseño e implementación por parte de un equipo de desarrolladores especializados en tecnologías web. Se realizaron pruebas exhaustivas para garantizar la fiabilidad y seguridad del sistema antes de su implementación final.</p> <p>Tras la implementación del sistema web, se observaran mejoras significativas en la productividad y eficiencia del equipo legal. La capacidad de acceder rápidamente a la información relevante y de colaborar de manera más efectiva contribuyó a una prestación de servicios más ágil y de mayor calidad para los clientes de la consultoría.</p> <p>En resumen, el caso de estudio destacó cómo la implementación de un sistema web personalizado puede transformar y mejorar la gestión de expedientes en una consultoría legal, optimizando los procesos internos y mejorando la experiencia del cliente.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.