Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,254

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,254 results for “pan”

Learn how ShareScore rates datasets ↗
zenodo40/100

Pan-European temperature distribution at depth - GeoDH project

<p>The dataset includes two shapefiles showing the temperature distribution at depth in Europe, specifically areas with temperatures exceeding 50&deg;C at 1000m depth and 90&deg;C at 2000m depth.<br><br>This dataset was developed for assessing the potential of Geothermal District Heating in Europe as part of the <strong>GeoDH project</strong> (<a href="http://geodh.eu/" target="_new" rel="noopener">http://geodh.eu/</a>). Please note that this represents the<strong> state of the art as of 2014</strong> and that geological, technological, and regulatory developments may have occurred since its creation, and users should verify if more recent data is available for their purposes.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Data for PAN at SemEval 2019 Task 4: Hyperpartisan News Detection

<p>Training, validation, and test data for the <a href="https://webis.de/events/semeval-19/">PAN @ SemEval 2019 Task 4: Hyperpartisan News Detection</a>.</p> <p>See the README for details.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

PAN Arabic Intrinsic Plagiarism Detection Shared Task Corpus

<p>Evaluation corpus for ARAbic INtrinsic plagiarism detection (InAra Corpus)&nbsp;</p> <p>&nbsp;</p> <p>This corpus has been used in AraPlagDet 2015 shared task&nbsp;</p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a>&nbsp;or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a>&nbsp;</p> <p>&nbsp;</p> <p><strong>I. SYNOPSIS&nbsp;</strong></p> <p>InAra corpus comprises 2048 documents; 80% of them contain passages&nbsp;borrowed from other documents to simulate documents that contain&nbsp;plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p>&nbsp;</p> <p><strong>II. DESCRIPTION&nbsp;</strong></p> <p>Each part of the corpus (training and test) consists mainly of 2 datasets:&nbsp;textual files and XML files.&nbsp;The textual files represent the suspicious documents i.e., the documents&nbsp;that contain artificial plagiarism; and the XML files are the plagiarism&nbsp;annotation i.e. they provide for each plagiarized passage its starting&nbsp;offset in the suspicious document and its length (offset and length are both expressed in characters). A suspicious document file and its plagiarism&nbsp;annotation file share the same name.</p> <p>&nbsp;</p> <p><strong>III. PURPOSE&nbsp;</strong></p> <p>The purpose of InAra corpus is to evaluate automatic plagiarism&nbsp;detection methods, notably methods of the intrinsic approach. This&nbsp;approach consists in uncovering the plagiarized passages on the basis of&nbsp;the writing style inconsistency in a given suspicious document. As&nbsp;opposed to the external approach, the intrinsic approach does not&nbsp;necessitate any comparison of the suspicious document against the&nbsp;potential sources of plagiarism. Hence, InAra corpus is not appropriate for the evaluation of the external plagiarism detection because the source of plagiarism are not provided.</p> <p>It should be noted that some documents in InAra corpus contain religious&nbsp;quotations (e.g., Quran and Hadith). These quotations have a peculiar writing style and then a simple intrinsic plagiarism detection software can consider them as plagiarism. However, quotations are not plagiarism, and they are not&nbsp;annotated in the XML files in InAra. Hence, it is an important feature for the plagiarism detection systems evaluated on InAra to not consider religious quotations as plagiarism cases unless they appear as part of a larger&nbsp; plagiarism case.</p> <p>&nbsp;</p> <p><strong>IV. BUILDING METHODS&nbsp;</strong></p> <p>The documents that compose InAra corpus do not contain actual plagiarism&nbsp;cases. They are rather artificial suspicious documents in which&nbsp;plagiarism was created automatically by a software that takes fragments&nbsp;of text from one or more sources documents and inserts them in another&nbsp;one according to a set of parameters, namely the percentage of plagiarism&nbsp;and the plagiarized passages lengths. This building method is the same&nbsp;used to construct PAN 2009-2011 corpora of plagiarism detection (see&nbsp;<a href="http://pan.webis.de">http://pan.webis.de</a> for more information on PAN competition and its&nbsp;corpora).&nbsp;</p> <p>&nbsp;</p> <p><strong>V. LANGUAGE AND ENCODING&nbsp;</strong></p> <p>All the textual documents of this corpus are written in Arabic language&nbsp;and encoded in UTF-8 without BOM.</p> <p>&nbsp;</p> <p><strong>VI. SOURCES OF TEXTS&nbsp;</strong></p> <p>Texts used to build this corpus, either suspicious documents or the&nbsp;inserted passages, are taken mainly from the open library Arabic&nbsp;Wikisource (http://ar.wikisource.org), one of Wikimedia Foundation&nbsp;projects. A few numbers of documents were taken from other websites,&nbsp;namely:&nbsp;</p> <ul> <li>Create your own country blog: http://diycountry.blogspot.com&nbsp;</li> <li>Corpus of Classical Arabic (KSUCCA): http://ksucorpus.ksu.edu.sa&nbsp;</li> <li>Islamic book web site: http://www.islamicbook.ws&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>VII. COPYRIGHT AND AVAILABILITY&nbsp;</strong></p> <p>We were very careful to build the corpus with copyright-free texts only,&nbsp;to be able to make it publicly available without any sort of problems&nbsp;with texts owners.&nbsp;</p> <p>&nbsp;</p> <p><strong>VIII. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using InAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on InAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>Additional information on the corpus building are in the papers:</p> <ul> <li>Bensalem, I., Rosso, P., Chikhi, S.: A New Corpus for the Evaluation&nbsp;of Arabic Intrinsic Plagiarism Detection. In: Forner, P., M&uuml;ller, H., Paredes, R., Rosso, P., and Stein, B. (eds.) CLEF 2013, LNCS, vol. 8138. pp. 53&ndash;58. Springer, Heidelberg (2013).</li> <li>Bensalem, I., Rosso, P., Chikhi, S.: Building Arabic Corpora from&nbsp;Wikisource. 10th ACS/IEEE International Conference on Computer Systems&nbsp;and Applications (AICCSA&rsquo;13),May 27-30 Fes/Ifran, Morocco (2013).IEEE.&nbsp;</li> </ul> <p>&nbsp;</p> <p>You may wish to&nbsp;compare the results of your experiments with the result of the following papers that used InAra corpus:</p> <ul> <li>Bensalem I, Rosso P, Chikhi S (2019) On the use of character n-grams&nbsp;as the only intrinsic evidence of plagiarism. Language Resources and&nbsp;Evaluation 53:363&ndash;396. doi: 10.1007/s10579-019-09444-w</li> <li>Mahgoub AY, Magooda A, Rashwan M, et al (2015) RDI System for&nbsp;Intrinsic Plagiarism Detection (RDI_RID), Working Notes for&nbsp;PAN-AraPlagDet at FIRE 2015. In: Majumder P, Mitra M, Agrawal M,&nbsp;&nbsp;Mehta P (eds) Post Proceedings of the Workshops at the 7th Forum for&nbsp;Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587. CEUR-WS.org, pp 129&ndash;130</li> </ul> <p>&nbsp;</p> <p><strong>IX. WARNING&nbsp;</strong></p> <p>It should be noted that the Arabic texts may contain quotations from the&nbsp;Quran and the Hadith; and due to the fact that text insertion is&nbsp;automatic and in random positions, it is possible that the plagiarized&nbsp;text is inserted unintentionally between Quranic verses or sentences of&nbsp;a Hadith cited in a document. Hence, the inserted passages may alter&nbsp;the meaning of the original text. For these reasons, this corpus must&nbsp;not be used outside the purpose for which it was built. Examples of the&nbsp;inappropriate use include using the corpus documents as a source of&nbsp;knowledge or distributing them without mentioning that they contain&nbsp;borrowed texts. If you are not interested in plagiarism detection and&nbsp;you are retaining the corpus because it contains books you want to read,&nbsp;then this corpus is not the right source. Please, you should refer to the&nbsp;</p> <p>sources mentioned in Section VI where you can find the original content of&nbsp;the books you are looking for. We emphasize that we are not responsible&nbsp;for the results of any use of this corpus other than the evaluation of&nbsp;the intrinsic plagiarism detection methods.&nbsp;</p> <p>&nbsp;</p> <p><strong>X. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using InAra&nbsp;corpus. Please do not hesitate to contact us with the following email&nbsp;address: bens.imene@gmail.com</p> <p>&nbsp;</p> <p>Imene Bensalem&sup1;, Paolo Rosso&sup2;, Salim Chikhi&sup1;</p> <p>&sup1;MISC Lab. Constantine 2 university, Algeria</p> <p>&sup2;PRHLT, Universitat Polit&egrave;cnica de Val&egrave;ncia, Spain&nbsp;</p>

opencc-by-4.0Jun 2015View details →
zenodo40/100

PAN Arabic External Plagiarism Detection Shared Task Corpus

<p>Evaluation Corpus for ARAbic EXternal plagiarism detection (ExAra Corpus)&nbsp;</p> <p>&nbsp;</p> <p>This corpus has been used in AraPlagDet 2015 shared task&nbsp;</p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a>&nbsp;or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a></p> <p>If you publish a paper about your experimentations using ExAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>&nbsp;</p> <p><strong>I. SYNOPSIS&nbsp;</strong></p> <p>ExAra corpus comprises 2345 documents; almost half of them (suspecious doucuments) contain passages borrowed from the other half (source docucments) to simulate documents that contain plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p>&nbsp;</p> <p><strong>II. DESCRIPTION&nbsp;</strong></p> <p>Each part of the corpus (training and test) consists mainly of 3 datasets:&nbsp; 2 sets of textual files and 1 set of XML files. The 2 sets of the textual files are the suspicious documents (i.e. the documents that contain artificial plagiarism) and the source documents (i.e., the documents from which the suspicious passages have been plagiarised). The 3rd set of documents contains XML files, which are the plagiarism annotation, i.e., they provide for each plagiarized passage its starting offset and its length in both the suspicious and source documents (offset and length were both expressed in characters). A suspicious document file (.txt) and its plagiarism annotation file (.xml) share the same name.</p> <p>&nbsp;</p> <p><strong>III. PURPOSE&nbsp;</strong></p> <p>The purpose of ExAra corpus is to evaluate automatic plagiarism detection methods, notably methods of the External approach. This approach consists in uncovering the plagiarized passages on the basis of their similarity with passages in the source documents.</p> <p>It should be noted that some suspicious documents in ExAra corpus contain religious quotations (e.g., Quran and Hadith) and common phrases. Some of them appear also in some source documents, and hence a simple plagiarism detection software can consider them as plagiarism. However, quotations and common phrases are legitimate text reuse cases and are not annotated in the XML files in ExAra. Therefore, it is an important feature for the&nbsp;plagiarism detection systems evaluated on ExAra to not consider religious quotations and common phrases as plagiarism cases unless they appear as part of a larger plagiarism case.</p> <p>&nbsp;</p> <p><strong>IV. BUILDING METHODS&nbsp;</strong></p> <p>The documents that compose ExAra corpus do not contain actual plagiarism&nbsp; cases, they are rather artificial suspicious documents in which&nbsp;plagiarism was created automatically by a software that takes fragments&nbsp;of text from one or more sources documents and inserts them in another&nbsp;one according to a set of parameters, namely the percentage of plagiarism&nbsp;and the lengths of the plagiarized passages. Some of the plagiarised fragments are&nbsp;obfuscated manually or automatically before inserting them in the suspicious documents.</p> <p>This building method is the same used to construct PAN 2009-2011 corpora of plagiarism detection (see http://pan.webis.de for more information on PAN competition and its&nbsp;corpora).&nbsp;</p> <p>&nbsp;</p> <p><strong>V. LANGUAGE AND ENCODING&nbsp;</strong></p> <p>All the textual documents of this corpus are written in Arabic language&nbsp;and encoded in UTF-8 without BOM.</p> <p>&nbsp;</p> <p><strong>VI. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using ExAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>&nbsp;</p> <p><strong>VII. WARNING&nbsp;</strong></p> <p>It should be noted that the Arabic texts may contain quotations from the&nbsp;Quran and the Hadith; and due to the fact that text insertion is&nbsp;automatic and in random positions, it is possible that the plagiarized&nbsp;text is inserted unintentionally between Quranic verses or sentences of&nbsp;a Hadith cited in a document. Hence, the inserted passages may alter&nbsp;the meaning of the original text. For these reasons, this corpus must&nbsp;not be used outside the purpose for which it was built. Examples of the&nbsp;inappropriate use include using the corpus documents as a source of&nbsp;knowledge or distributing them without mentioning that they contain&nbsp;borrowed texts. If you are not interested in plagiarism detection, and you are retaining the corpus because it contains articles you want to read,&nbsp;then this corpus is not the right source. Please, you should refer to the&nbsp;sources mentioned in (Bensalem et al. 2015) (i.e.,the paper above) where&nbsp;you can find the original content of the articles you are looking for.</p> <p>We emphasize that we are not responsible for the results of any use of this corpus other than the evaluation of the external plagiarism detection methods.&nbsp;</p> <p>&nbsp;</p> <p><strong>VIII. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using ExAra&nbsp;corpus. Please do not hesitate to contact us with the following email&nbsp;address: bens.imene@gmail.com</p> <p>&nbsp;</p> <p>Imene Bensalem&sup1;, Imene Boukhalfa&sup1;, Paolo Rosso&sup2;, Salim Chikhi&sup1;</p> <p>&sup1;MISC Lab. Constantine 2 university, Algeria</p> <p>&sup2;PRHLT, Universitat Polit&egrave;cnica de Val&egrave;ncia, Spain&nbsp;</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

The pan-genome unearths gene content and transposable element variations in modern pigs

<p>Genes, gene annotations, proteins, and sequences identified in the non-reference genome of the pig pan-genome. Transposable insertion polymorphisms (TIP) indentified in the pig mobolome.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Text-fig. 10. Extant pans east of Inhaminga (18°26′28″'S: 35°35′45″E) surrounded by woodland. The pans typically have an arid, vegetation-free, marginal zone and a water-logged sump. Some pans are connected to each other by shallow overflow valleys. Image modified from Google Earth. in Stratigraphy, Chronology And Palaeontology Of The Tertiary Rocks Of The Cheringoma Plateau, Mozambique

Text-fig. 10. Extant pans east of Inhaminga (18°26′28″'S: 35°35′45″E) surrounded by woodland. The pans typically have an arid, vegetation-free, marginal zone and a water-logged sump. Some pans are connected to each other by shallow overflow valleys. Image modified from Google Earth.

opencc-by-4.0Dec 2021View details →
zenodo40/100

ENTSO-E Pan-European Climatic Database (PECD 2021.3) in Parquet format

<p><strong>ENTSO-E Pan-European Climatic Database (PECD 2021.3) in Parquet format</strong></p> <p><strong>TL;DR</strong>: this is a tidy and friendly version of a subset of the PECD 2021.3 data by ENTSO-E: hourly capacity factors for wind onshore, offshore, solar PV, hourly electricity demand, weekly inflow for reservoir and pumping and daily generation for run-of-river. All the data is provided for &gt;30&nbsp;climatic years (1982-2019 for wind and solar, 1982-2016 for demand, 1982-2017 for hydropower)&nbsp;and at national and sub-national (&gt;140 zones) level.</p> <p><strong>UPDATE (19/10/2022):&nbsp;</strong>updated the demand files due after fixing a bug in the processing code (the file for 2030 was the same for 2025) and solving an issue caused by a malformed header in the ENTSO-E excel files.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>ENTSO-E has released with the latest European Resource Adequacy Assessment (<a href="https://www.entsoe.eu/outlooks/eraa/">ERAA 2021</a>) all the inputs used in the study.<br> Those inputs include:<br> - Demand dataset: <a href="https://eepublicdownloads.azureedge.net/clean-documents/sdc-documents/ERAA/Demand%20Dataset.7z">https://eepublicdownloads.azureedge.net/clean-documents/sdc-documents/ERAA/Demand%20Dataset.7z</a><br> - Climate data: <a href="https://eepublicdownloads.entsoe.eu/clean-documents/sdc-documents/ERAA/Climate%20Data.7z">https://eepublicdownloads.entsoe.eu/clean-documents/sdc-documents/ERAA/Climate%20Data.7z</a></p> <p>The data files and the methodology are available on the <a href="https://www.entsoe.eu/outlooks/eraa/2021/eraa-downloads/">official webpage</a>.&nbsp;</p> <p>As done for the previous releases (see <a href="https://zenodo.org/record/3702418#.YbmhR23MKMo">https://zenodo.org/record/3702418#.YbmhR23MKMo</a> and <a href="https://zenodo.org/record/3985078#.Ybmhem3MKMo">https://zenodo.org/record/3985078#.Ybmhem3MKMo</a>), the original data - stored in large Excel spreadsheets - have been tidied and formatted in open and friendly formats (CSV for the small tables and Parquet for the large files)</p> <p>Furthermore, we have carried out a simple country-aggregation for the original data - that uses instead &gt;140 zones.</p> <p><strong>DISCLAIMER</strong>: <em>the content of this dataset has been created with the greatest possible care. However, we invite to use the original data for critical applications and studies.&nbsp;</em></p> <p><strong>Description</strong></p> <p>This dataset includes the following files:</p> <p>- <em>capacities-national-estimates.csv</em>: installed capacity in MW per zone, technology and the two scenarios (2025 and 2030). The files include also the total capacity for each technology per country (sum of all the zones within a country)<br> - <em>PECD-2021.3-wide-LFSolarPV-2025</em> and <em>PECD-2021.3-wide-LFSolarPV-2030</em>: tables in Parquet format storing in each row the capacity factor for solar PV for a hour of the year and all the climatic years (1982-2019) for a specific zone. The two files contain the capacity factors for the scenarios &quot;National Estimates 2025&quot; and &quot;National Estimates 2030&quot;<br> - <em>PECD-2021.3-wide-Onshore-2025</em> and<em> PECD-2021.3-wide-Onshore-2030</em>: same as above but for wind onshore<br> - <em>PECD-2021.3-wide-Offshore-2025</em> and <em>PECD-2021.3-wide-Offshore-2030</em>: same as above but for wind offshore<br> - <em>PECD-wide-demand_national_estimates-2025</em> and<em> PECD-wide-demand_national_estimates-2030</em>: hourly electricity demand for all the climatic years for a specific zone. The two files contain the load for the scenarios &quot;National Estimates 2025&quot; and &quot;National Estimates 2030&quot;&nbsp;<br> - <em>PECD-2021.3-country-LFSolarPV-2025</em> and <em>PECD-2021.3-country-LFSolarPV-2030</em>: tables in Parquet format storing in each row the capacity factor for country/climatic year and hour of the year. The two files contain the capacity factors for the scenarios &quot;National Estimates 2025&quot; and &quot;National Estimates 2030&quot;<br> -<em> PECD-2021.3-country-Onshore-2025</em> and<em> PECD-2021.3-country-Onshore-2030</em>: same as above but for wind onshore<br> -<em> PECD-2021.3-country-Offshore-2025</em> and <em>PECD-2021.3-country-Offshore-2030</em>: same as above but for wind offshore<br> - <em>PECD-country-demand_national_estimates-2025</em> and <em>PECD-country-demand_national_estimates-2030</em>: same as above but for electricity demand<br> - <em>PECD_EERA2021_reservoir_pumping.zip</em>: archive with four files per each scenario: 1. table.csv with generation and storage capacities per zone/technology, 2. zone weekly inflow (GWh), 3. table.csv with generation and storage per country/technology and 4. country weekly inflow (GWh)<br> - <em>PECD_EERA2021_ROR.zip</em>: as for the previous file but the inflow is daily<br> - <em>plots.zip</em>: archive with 182 png figures with the weekly climatology for all the variables (daily for the electricity demand)</p> <p><strong>Note</strong></p> <p>I would like to thank Laurens Stoop for sharing the onshore wind data for the scenario 2030, that was corrupted in the original archive.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

CONSOLE_WP3_Task3.2_ Pan-EU survey of farmers and other rural landowners_AT_2022.10.25

<p>Dataset containing information about the structural equation model carried out in Austria for the CONSOLE (CONtract Solutions for Effective and lasting delivery of agri-environmental-climate public goods by EU agriculture and forestry) Project</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

The Pan-Canadian Chemical Library: A Mechanism to Open Academic Chemistry to High-Throughput Virtual Screening

<h1>Pan-Canadian Chemical Library</h1> <p>This Zenodo repository contains the cheap and druglike subset of the Pan-Canadian Chemical Library (PCCL) project. For more information, visit&nbsp;<a href="https://pccl.thesgc.org/" rel="nofollow">https://pccl.thesgc.org</a>.</p> <h2>PCCL library</h2> <p>The PCCL library is splitted by reaction, then by number of heavy atoms. Two types of files are available in zip archives:</p> <ul> <li>The SMILES format files, with the SMILES string and their product name,</li> <li>The CSV format file, with all the information generated during their enumeration: reagents, druglike properties, etc.</li> </ul> <p>Note: Purchasability is defined according to two integers: 1 for products only composed of BB-50 reagents, and 2 for products composed of BB-40 or BB-50 reagents. Read more about the meaning of these reagents groups in the article below.</p> <h2>Citation</h2> <p>If you find the PCCL useful or if you use it, please cite our paper:</p> <p>Bedart, C. <em>et al.</em> The Pan-Canadian Chemical Library: A mechanism to open academic chemistry to high-throughput virtual screening. Scientific Data 11, (2024).<br>doi: <a title="10.1038/s41597-024-03443-5" href="https://www.nature.com/articles/s41597-024-03443-5">10.1038/s41597-024-03443-5</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Fig. 7 in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 7. SEM images of the jaw symphyseal region and details of tooth crowns in two extant gekkotans, Gonatodes albogularis (Duméril and Bibron, 1836) A) and Sphaerodactylus klauberi Grant, 1923 (B), showing a similar sulcus on the crown as present in Bauersaurus gen. nov.

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 4 in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 4. Pan-gekkotan squamate Bauersaurus cosensis gen. et sp. nov. from Cos, France, early Eocene (MP 10–11). Tooth details of the holotype (UM- COS-1012) maxilla (A) and UM-COS-1013 dentary (B) in ventromedial (A1), ventral (A2), anteromedial (A3), dorsal (B1), anterior (B2) views (all micro-CT visualizations).

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 6 in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 6. Maxillae of several gekkotans compared with Bauersaurus gen. nov. in lateral (left column) and medial (right column) views; 3D models of gekkotans were mirrored. Top to bottom, Sphaerodactylus grandisquamis spanius Stejneger, 1904, RT 14708 (Sphaerodactylidae), Holodactylus africanus Boettger, 1893, CAS 198932 (Eublepharidae), Paradelma orientalis (Günther, 1876) CAS 77652 (Pygopodide), Underwoodisaurus milii Bory de Saint-Vincent, 1823, CAS 74744 (Carphodactylidae), Pseudothecadactylus australis (Günther, 1877) MCZ Herp R- 35162 (Diplodactylide), Bauersaurus cosensis the holotype UM-COS-1012 (Pan-Gekkota). Bauersaurus cosensis is inferred as sister to Gekkota.

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 1 in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 1. Pan-gekkotan squamate Bauersaurus cosensis gen. et sp. nov. from Cos, France, early Eocene (MP 10–11). Holotype (UM-COS-1012), right maxilla, in lateral (A1), medial (A2), dorsal (A3), ventral (A4), anteromedial (A5), and dorsomedioposterior (A6) views (all micro-CT visualizations).

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 3 in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 3. Pan-gekkotan squamate Bauersaurus cosensis gen. et sp. nov. from Cos, France, early Eocene (MP 10–11). Photographs of the holotype (UM- COS-1012) maxilla (A) and UM-COS-1013 dentary (B) in lateral (A1, B1) and medial (A2, B2) views.

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 5. A in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 5. A. Reconstruction of the complete skull of an early Eocene pan-gekkotan squamate Bauersaurus cosensis gen. et sp. nov. from Cos, France, using a D model of a living gekkotan with similar morphology and proportions of bones, Pseudothecadactylus australis (MCZ Herp R- 35162), grey silhouette; Bauersaurus, micro-CT visualizations of jaws in lateral view. B–F. Comparison of the overall shape of the facial process with other gekkotans, i.e. the Late Jurassic Eichstaettisaurus schroederi (the holotype BSPG 1937 I; modified from Simões et al. 2017), (B) the early Eocene (MP 10) Laonogekko lefevrei (the holotype MNHN PMT 5; modified from Augé 2005) (C), the middle Eocene (MP 12 and MP 13) Geiseleptes delfinoi (the holotype GMH Ce IV-4057-1933; modified from Villa et al. (2022) (D), the middle–late Eocene (MP 16–19) Cadurcogekko piveteaui (the right maxilla MNHN QU 17734; modified from Augé 2005) (E), and Cadurcogekko cf. piveteaui (mirror inverted left maxilla NHMW 2019/0052/0001 originally described by Georgalis et al. 2021) (F).

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 2 in An early Eocene pan-gekkotan from France could represent an extra squamate group that survived the K/Pg extinction

Fig. 2. Pan-gekkotan squamate Bauersaurus cosensis gen. et sp. nov. from Cos, France, early Eocene (MP 10–11). UM-COS-1013, right dentary, in lateral (A1), medial (A2), dorsal (A3) views (all micro-CT visualizations).

opencc-by-4.0Nov 2023View details →
zenodo40/100

Fig. 14. A in Morphology and relationships of the enigmatic stenothecoid pan-brachiopod Stenothecoides-new data from the middle Cambrian Burgess Shale Formation

Fig. 14. A. Hypothesized morphogenetic pathway for derivation of the bivalved stenothecoid scleritome from a multisclerite, tubular, organocalcitic, eccentrothecimorph ancestor via a hypothetical organocalcitic tannuolinid. Eccentrothecimorph slightly modified from Skovsted et al. (2011: text-fig. 17). Tannuolinid modified from Skovsted et al. (2014). B. A possible cladogram that assumes an organocalcitic scleritome is primitive for the Pan-Brachiopoda. In this view, organophosphatic mineralogy evolved independently several times and the Stenothecoida evolved a bivalved scleritome independently of the Brachiopoda + Micrina clade.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Fig. 7 in Morphology and relationships of the enigmatic stenothecoid pan-brachiopod Stenothecoides-new data from the middle Cambrian Burgess Shale Formation

Fig. 7. SEM micrographs of stenothecoid pan-brachiopod Stenothecoides cf. elongata, middle Cambrian, Burgess Shale Formation, Kootenay National Park, Canada, Locality 3. A. ROMIP 66248, articulated shell in right lateral view. B. ROMIP 66249, articulated shell in left lateral (B1) and posterior (B2)

opencc-by-4.0Dec 2021View details →
zenodo40/100

Fig. 5 in Morphology and relationships of the enigmatic stenothecoid pan-brachiopod Stenothecoides-new data from the middle Cambrian Burgess Shale Formation

Fig. 5. Stenothecoid pan-brachiopod Stenothecoides rasettii sp. nov., middle Cambrian, Burgess Shale Formation, Yoho National Park, Canada, Locality 2 (A, G) and Locality 1 (B–F, H–J). A. TMP 2002.083.0178, dorsal valve in interior (A1), exterior (A2), right lateral (A3), and posterior (A4) views. B. TMP 2008.024.1138, ventral valve in interior view. C. TMP 2008.024.1133, dorsal valve in interior view, arrows show possible bifurcated peripheral ridges. D. TMP 2008.024.1143, dorsal valve in interior (D1) and anterior oblique (D2) views. E. TMP 2008.024.1144, dorsal valve in interior (E1) and anterior (E2) views. F. TMP 2008.024.1134, ventral valve, interior view. G. TMP 2002.083.0177 (same specimen as Fig. 4B), ventral valve in interior (G1, magnified) and oblique (G2) views. H. TMP 2008.024.1135, dorsal valve in interior view. I. TMP 2008.024.1121, ventral valve, anterior view. J. TMP 2008.024.1136, ventral valve in interior view, showing detached apical boss and remnant apical stem and cardinal troughs.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Fig. 4 in Morphology and relationships of the enigmatic stenothecoid pan-brachiopod Stenothecoides-new data from the middle Cambrian Burgess Shale Formation

Fig. 4. Stenothecoid pan-brachiopod Stenothecoides rasettii sp. nov., middle Cambrian, Burgess Shale Formation, Yoho National Park, Canada, Locality 2 A, B, D) and Locality 1 (C, D). A. TMP 2002.083.0176 (holotype), dorsal valve in exterior (A1) and interior (A2) views. B. TMP 2002.083.0177, ventral valve in exterior view. C. TMP 2008.024.1147, dorsal valve, internal apical area. D. TMP 2008.024.1122, articulated juvenile, dorsal/ventral (uncertain) valve in posterior (D1), lateral (D2), and oblique planar (D3) views; D1 and D3 show posterior median opening, D2 shows slightly sinusoidal commissure.

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record