Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10,812

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10,812 results for “novels”

Learn how ShareScore rates datasets ↗
zenodo48/100

Novel estimates of the leaf relative uptake rate of carbonyl sulfide from optimality theory

<p>Data and Matlab scripts for repeating the analysis presented in the paper. In addition, global monthly climatological LRUs are provided at 0.05&deg; resolution for the period 2001-2010 as nc-files.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

A novel and holistic approach for experimental X-ray fundamental parameter determination - the Ru L-shell

<p>This dataset contains the experimentally determined fundamental parameters for the ruthenium L-subshells from our paper with the title &quot;A novel and holistic approach for experimental X-ray fundamental parameter determination - the Ru L-shell&quot;. The paper will be published soon in a peer-reviewd journal.</p> <p>This file contains L-subshell fluorescence yields, L-shell Coster-Kronig factors, L-shell Auger yields,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br> Mass attenuation coefficients in energy range from 2.41 keV to 8 keV, L-subshell photo ionization cross sections up to 8 keV and&nbsp;<br> L-subshell fluorescence prodution cross sections of Ru.</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Dataset for Evaluation of a novel microfluidic chip-like device for purifying bovine frozen-thawed semen for in vitro fertilization

<p>VetCount<sup>TM</sup> Harvester (MotilityCount ApS, Copenhagen, Denmark) is a novel sperm<br> purification device. It consists of two chambers separated by a 10 &mu;M microporous<br> membrane. Untreated semen is applied in one chamber, sperm collection medium<br> in the other. Motile sperm cells are selected by actively swimming through the<br> membrane pores into the medium containing chamber. After 30 min incubation,<br> the sperm collection medium can be aspirated and the purified sperm is ready<br> for further use.<br> In a first experiment we assessed sperm quality and recovery of frozen-thawed semen<br> from six different bulls (n = 6) prior to and after purification with the<br> VetCount<sup>TM</sup> Harvester or BoviPure<sup>TM</sup> gradient centrifugation,&nbsp;a commercial available&nbsp;standard technique. In a second approach, a competitive fertilization assay was performed. Ten straws per bull were pooled, split<br> in two subsamples, and simultaneously purified either with the VetCount<sup>TM</sup> Harvester<br> or BoviPure<sup>TM</sup> gradient centrifugation. Following purification, sperm cells<br> from each treatment group were fluorescently labeled with either MitoTracker<sup>TM</sup><br> Red FM or MitoTracker<sup>TM</sup> Green FM. <em>In vitro</em> matured oocytes were inseminated&nbsp;with<br> equal numbers of red and green labeled sperm. Eighteen hours after fertilization,<br> fluorescence microscopy was used to determine the origin of the fertilizing spermatozoon.</p>

opencc-by-4.0Jun 2023View details →
zenodo48/100

Data publication supplementing "Novel nanoindentation strain rate sweep method for continuously investigating the strain rate sensitivity of materials at the nanoscale"

<p>This data publication contains the results of nanoindentation tests on Fused silica, nanocrystalline nickel, a nanocrystalline FeCr alloy, a bulk metallic glass, the superplastic alloy Zn-22%Al and single crystalline aluminum as well as the method files developed for the G200 nanoindenter. It supplements the publication "Novel nanoindentation strain rate sweep method for continuously investigating the strain rate sensitivity of materials at the nanoscale". The materials are described in more detail in the respective publication. The data publication takes over the sample naming convention from the related publication.&nbsp;</p><p>Nanoindentation measurement were performed by H. Holz at the Max-Planck Institut für Eisenforschung GmbH, Max-Planck-Straße 1, 40237 Düsseldorf, Germany using a G200 nanoindenter (KLA, Milpitas, CA, USA), equipped with a modified Berkovich diamond indenter tip of the type 171-561-500 with a serial number of C-0040446 from Synton MDP (Nidau, Switzerland). Constant strain rate tests, strain rate jump tests, strain rate sweep tests and strain rate sweep reversal tests were performed on each material in Continuous Stiffness Measurement (CSM) mode. The maximum indentation depth was 2200 nm, the CSM amplitude 2 nm and the CSM frequency 45 Hz. Strain rates were varied within the range 0.001 – 0.1 s-1. Further information on the test protocol can be found in the corresponding publication.</p><p>The subfolder "Nanoindentation data" contains the raw data for all valid indents as output by the NanoSuite © software v 7.1.7 and converted to the semicolon-separated format. The naming convention for the folder in which the CSV files are located in gives first the used material, then the method used with additional information to the parameters inputted for the method such as strain rate and indentation depth all separated by an underscore. An example can be "FS_CSR_01s-1" for a constant strain rate tests performed on fused silica with a strain rate target of 0.1 s-1 or "Nc-Ni_SweepReversal_005s-1_0005s-1" for a sweep reversal test performed on the nanocrystalline nickel sample with a targeted initial and ending strain rate of 0.05 s-1 and a strain rate target at which the strain rate direction gets reversed of 0.005 s-1. The CSV files are named either "Results" giving the average results of each test, "Required Inputs" giving information about the parameters used for the experiments, "Inputs Editable Post Test" giving information about the analysis parameters to obtain the results from, and "Test XXX" which include the Raw data of the corresponding test number. In each file the first row gives the data description e.g., "Time", the second row the physical unit e.g., "s" for seconds and from the third row the measured values.</p><p>The Nano Suite method files to perform the experiments on KLA G200 instruments is provided in the folder "G200 methods". The method for the strain rate sweep experiments is called "Strain Rate Sweep.msm" and the method for the strain rate sweep reversal experiments "Strain Rate Sweep Reversal.msm". This method is provided as is and shall be used at your own risk. The authors explicitly decline responsibility for any physical or immaterial damage resulting from the use of this method. Should minor issues occur, some feedback to the authors would be greatly appreciated.</p>

opencc-by-4.0Oct 2023View details →
edi48/100

Fire-severity effects on plant-fungal interactions after a novel tundra wildfire disturbance: implications for arctic shrub and tree migration

Background-Vegetation change in high latitude tundra ecosystems is expected to accelerate due to increased wildfire activity. High-severity fires increase the availability of mineral soil seedbeds, which facilitates recruitment, yet fire also alters soil microbial composition, which could significantly impact seedling establishment. Results - We investigated the effects of fire severity on soil biota and associated effects on plant performance for two plant species predicted to expand into Arctic tundra. We inoculated seedlings in a growth chamber experiment with soils collected from the largest tundra fire recorded in the Arctic and used molecular tools to characterize root-associated fungal communities. Seedling biomass was significantly related to the composition of fungal inoculum. Biomass decreased as fire severity increased and the proportion of pathogenic fungi increased. Conclusions - Our results suggest that effects of fire severity on soil biota reduces seedling performance and thus we hypothesize that in certain ecological contexts fire-severity effects on plant-fungal interactions may dampen the expected increases in tree and shrub establishment after tundra fire.

openOpenMar 2016View details →
zenodo44/100

Data and R Code from "A novel approach to sustainability assessment of food supply chains using networks of ecosystem services"

<p>Data and R code from this paper applying network analysis (iGraph) to&nbsp;two case studies pre and post agroecological transitions in Central America and Tanzania, Africa from the IPES-Food report. Further descriptions of this data and code can be found within the extended manuscript. R Code relies on the data from the scenarios (e.g., Nodes and Relations CSVs) and creates the output network metrics (e.g., Node Metric CSVs).&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Trophic cascade driven by behavioural fine-tuning as naïve prey rapidly adjust to a novel predator

<p>The arrival of novel predators can trigger trophic cascades driven by shifts in prey numbers. Predators also elicit behavioural change in prey populations, via phenotypic plasticity and/or rapid evolution, and such changes may also contribute to trophic cascades. Here we document rapid demographic and behavioural changes in populations of a prey species (grassland melomys <em>Melomys burtoni</em>, a granivorous rodent) following the introduction of a novel marsupial predator (northern quoll <em>Dasyurus hallucatus</em>). Within months of quolls appearing, populations of melomys exhibited reduced survival and population declines relative to control populations. Quoll-invaded populations (<em>n </em>= 4) were also significantly shyer than nearby, quoll-free populations (<em>n </em>= 3) of conspecifics. This rapid but generalised response to a novel threat was replaced over the following two years with more threat-specific antipredator behaviours (i.e. predator-scent aversion). Predator-exposed populations, however, remained more neophobic than predator-free populations throughout the study. These behavioural responses manifested rapidly in changed rates of seed predation by melomys across treatments. Quoll-invaded melomys populations exhibited lower per-capita seed take rates, and rapidly developed an&nbsp;avoidance of seeds associated with quoll scent, with discrimination playing out over a spatial scale of tens of metres. Presumably the significant and novel predation pressure induced by quolls drove melomys populations to fine-tune behavioural responses to be more predator-specific through time. These behavioural shifts could reflect individual plasticity (phenotypic flexibility) in behaviour or may be adaptive shifts from natural selection imposed by quoll predation. Our study provides a rare insight into the rapid ecological and behavioural shifts enacted by prey to mitigate the impacts of a novel predator and shows that trophic cascades can be strongly influenced by behavioural as well as numerical responses.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

Scored protein-protein interactions accompanying "A pan-plant protein complex map reveals deep conservation and novel assemblies"

<p><a href="http://plants.proteincomplexes.org/static/data/panplant_cfms_scores_annot.txt.gz">All scored pairwise protein-protein interactions with CF-MS scores (3,076,999 unique pairwise interactions)</a></p> <ul> <li>Description: Scores between Orthogroups with the corresponding CF-MS score and eggNOG generated orthogroup descriptions.</li> <li>Note: Only the highest scoring pairs are considered significant. A CF-MS score &gt;= 0.509 corresponds to 10% FDR, &gt;= 0.207 corresponds to 50% FDR</li> <li>Format: OrthogroupID1 [tab] OrthogroupID2 [tab] Score [tab] Annotation1 [tab] Annotation2</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Enhanced genome annotation strategy provides novel insights on the phylogeny of 'Flaviviridae': Supplementary material

<p>SUPPLEMENTARY MATERIAL</p> <p><strong>Index</strong></p> <ul> <li> <p>Table S1 (tableS1.csv): genomic data.</p> </li> <li> <p>Table S2 (tableS2.csv): character categorization for selected nodes.</p> </li> <li> <p>Table S3 (tableS3.csv): programs and parameters.</p> </li> <li> <p>Table S4 (tableS4.csv): annotation efficiency.</p> </li> <li> <p>File S1 (fileS1.gff): gene annotation.</p> </li> <li> <p>File S2 (fileS2.xml): configuration file for BEAST 2 (configuration.xml).</p> </li> <li> <p>Figure S1 (figureS1.pdf): dendrogram depicting the hierarchical clusters of trees based on match-split distances.</p> </li> <li> <p>Figure S2 (figureS2.pdf): full version of the working phylogenetic hypothesis (tree No. 0 in table 1).</p> </li> </ul> <p><strong>Figure captions</strong></p> <ul> <li>Figure S1: A dendrogram depicting the hierarchical clusters of trees based on match-split distances.&nbsp;Outgroup sequences (<em>Hepacivirus</em>, <em>Pegivirus</em>, and <em>Pestivirus</em>) were removed to guarantee the compared tree topologies would have the same terminals. Tree numbers correspond to those in table 1 of the manuscript. I. No outgroup sequences; some matrices were partitioned. II. Outgroup sequences and partitioned matrices. *This tree was produced without outgroup sequences.</li> <li>Figure S2: Full version of the working phylogenetic hypothesis (tree No. 0 in table 1). Branch lengths represent an estimation of the number of substitutions per site. Node labels indicate SH-aLRT support / ultrafast bootstrap (only shown if one of there is a value&nbsp;below 90%). Clade names correlate to the character categorization analysis (see table S2). Branch labels represent the four genera: I = <em>Pestivirus</em>; II = <em>Pegivirus</em>; III = <em>Hepacivirus</em>; IV = <em>Flavivirus</em>. *&nbsp;The Ecuador Paraiso Escondido virus (EPEV) was isolated from sand flies (<em>Psathyromyia abonnenci</em>). The EPEV was the first sand fly-borne <em>Flavivirus</em> identified in the New World.</li> </ul> <p><strong>Manuscript title</strong></p> <p>FLAVi: an enhanced annotator for viral genomes of <em>Flaviviridae</em>.</p> <p><strong>Authors</strong></p> <ul> <li> <p>de Bernardi Schneider, Adriano. University of California San Diego. ORCID: 0000-0001-7487-266X.</p> </li> <li> <p>Jacob Machado, Denis. University of North Carolina at Charlotte. ORCID: 0000-0001-9858-4515. Corresponding author.</p> </li> <li> <p>Guirales, Sayal.&nbsp;University of North Carolina at Charlotte.</p> </li> <li> <p>Janies, Daniel. University of North Carolina at Charlotte.</p> </li> </ul> <p><em>First author</em>: Adriano de Bernardi Schneider and Denis Jacob Machado have contributed equally to the manuscript.</p> <p><strong>Contact information</strong></p> <ul> <li> <p>Corresponding author: Denis Jacob Machado, Ph.D.</p> </li> </ul> <ul> <li> <p>OrcID: 0000-0001-9858-4515.</p> </li> </ul> <ul> <li> <p>Email: dmachado [at] uncc.edu.</p> </li> </ul> <p><strong>Other additional material</strong></p> <ul> <li>In addition to the material listed above, all 31 tree topologies and 15 alignment matrices discussed in this manuscript will are available in TreeBASE (<a href="http://purl.org/phylo/treebase/phylows/study/TB2:S24096">http://purl.org/phylo/treebase/phylows/study/TB2:S24096</a>) after the publication of the manuscript.</li> <li>The FLAVi pipeline and all the original scripts are available at GitLab (<a href="https://gitlab.com/MachadoDJ/FLAVi">https://gitlab.com/MachadoDJ/FLAVi</a>).</li> <li>The web application can be accessed at <a href="http://flavi-web.com">http://flavi-web.com</a>.</li> </ul>

opencc-by-4.0Jun 2019View details →
zenodo44/100

COVID-19 Tweets : A dataset contaning more than 600k tweets on the novel CoronaVirus

<p>This dataset contains&nbsp;653 996&nbsp;tweets related to the Coronavirus topic and highlighted by hashtags such&nbsp;as: #COVID-19, #COVID19, #COVID, #Coronavirus, #NCoV and #Corona. The tweets&#39; crawling period started on the 27<sup>th</sup> of February and ended on the 25<sup>th</sup> of March 2020, which is spread over four weeks.&nbsp;</p> <p>The tweets were generated by 390 458 users from 133 different countries and were written in 61 languages. English being the most used language with almost 400k tweets, followed by Spanish with around 80k tweets.&nbsp;</p> <p>The data is stored in as a CSV file, where each line represents a tweet. The CSV file provides information on the following fields:</p> <ul> <li>Author: the user who posted the tweet</li> <li>Recipient: contains the name of the user in case of a reply, otherwise it would have the same value as the previous field</li> <li>Tweet: the full content of the tweet</li> <li>Hashtags: the list of hashtags present in the tweet</li> <li>Language: the language of the tweet</li> <li>Relationship: gives information on the type of the tweet, whether it is a retweet, a reply, a tweet with a mention, etc.&nbsp;</li> <li>Location: the country of the author of the tweet, which is unfortunately not always available</li> <li>Date: the publication date of the tweet</li> <li>Source: the device or platform used to send the tweet</li> </ul> <p>The dataset can as well be used to construct a social graph since it includes the relations &quot;Replies to&quot;, &quot;Retweet&quot;, &quot;MentionsInRetweet&quot; and&nbsp;&quot;Mentions&quot;.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

CMU-MisCov19: A Novel Twitter Dataset for Characterizing COVID-19 Misinformation

<p>From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. Detection and characterization of misinformation requires an availability of annotated datasets. Most of the published COVID-19 Twitter datasets are generic, lack annotations or labels, employ automated annotations using transfer learning or semi-supervised methods, or are not specifically designed for misinformation. Annotated datasets are either only focused on &quot;fake news&quot;, are small in size, or have less diversity in terms of classes.</p> <p>Here, we present a novel Twitter misinformation dataset called <strong>&quot;CMU-MisCov19&quot;</strong> with 4573 annotated tweets over 17 themes around the COVID-19 discourse.&nbsp;We also present our annotation codebook for the different COVID-19 themes&nbsp;on Twitter, along with their descriptions and examples,&nbsp;for the community to use for collecting further annotations. Further details related to the dataset, and our analysis based on this dataset can be found at&nbsp;<a href="https://arxiv.org/abs/2008.00791">https://arxiv.org/abs/2008.00791</a>. In adherence to the Twitter&rsquo;s terms and conditions, we&nbsp;do not provide&nbsp;the full tweet JSONs but provide a &quot;.csv&quot; file with the tweet IDs so that the tweets&nbsp;can be rehydrated. We also provide the annotations, and the date of creation for each tweet for the reproduction of the results of our analyses.</p> <p><strong>Note: If for any reason, you are not able to rehydrate all the tweets, reach out to&nbsp;Shahan Ali Memon at (shahan@nyu.edu).</strong></p> <p>If you use this data, please cite our paper as follows:&nbsp;</p> <p><em>&quot;Shahan Ali Memon and Kathleen M. Carley. Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset, In Proceedings of The 5th International Workshop on Mining Actionable Insights from Social Networks (MAISoN 2020), co-located with CIKM, virtual event due to COVID-19, 2020.&quot;</em></p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Simulated NGS datasets for real-time detection of novel pathogens

<p>Datasets based on the <a href="https://doi.org/10.5281/zenodo.3678563">bacterial</a> and <a href="https://doi.org/10.5281/zenodo.4312525">viral</a> simulated NGS datasets. Fastq files correspond to tests sets of those datasets. Basecall files were generated based on the fastq files with an 8nt simulated barcode between the mates of a read pair. The &quot;rn&quot; datasets containg random length subreads (25-250bp) of the original validation and training reads.</p> <p>The Nanopore datasets were resimulated with <a href="https://github.com/liyu95/DeepSimulator">DeepSimulator 1.5</a> (Li et al., 2020) based on the original datasets (i.e. using the same species composition as the original data). The test Nanopore dataset contains full reads (target average length: 8kb) and the training and validation datasets - 250bp subreads.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

III PhasAGE International Conference - Design of novel functional amyloid assemblies - Lecture

<p>The&nbsp;III PhasAGE International Conference&nbsp;"Multiscale understanding of protein aggregation and biomolecular condensates in aging and disease" brought together members of the PhasAGE consortium as well as outstanding international speakers from multidisciplinary fields dedicated to unraveling the intricacies of protein aggregation and biomolecular condensates in the context of aging and disease. For details on the conference program please see&nbsp;https://phasage.eu/iii-phasage-international-conference/.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A scoping review on bovine tuberculosis highlights the need for novel data streams and analytical approaches to curb zoonotic diseases

<p>The following data and scripts are part of the manuscript titled 'A scoping review on bovine tuberculosis highlights the need for novel data streams and analytical approaches to curb zoonotic diseases' which is currently going through the peer-review process and has already been published as a preprint. Please read the README.txt file for information on the files uploaded.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A novel educational approach for safe endodontic syringe irrigation: a randomized controlled study

<p><span>(1) Educational video emphasizing the fundamentals of safe irrigation practices, incorporating evidence-based guidelines on appropriate plunger forces and the required time for safe irrigant delivery.&nbsp;</span></p> <p><span>(2) Dataset comprising&nbsp;the measurement data collected and processed during this study.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

GWAS summary stats in "Genome-wide association meta-analysis identifies two novel loci associated with dental caries."

<p>Summary stats of the genome-wide meta-analysis for dental caries and periodontal diseases in our study (population A and B).</p> <p>Article "Genome-wide association meta-analysis identifies two novel loci associated with dental caries."</p> <p>https://doi.org/10.1186/s12903-024-04799-1<br><br></p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Novel Libraries in Stack Overflow Posts

<p># Summary</p> <p>We present datasets detailing the appearance of novel libraries and library pairs in Stack Overflow posts in 12 languages between 2008 and 2023.</p> <div> <div># Disclaimer</div> <br> <div>Pair of libraries are displayed in the canonical format of &lt;lib_a&gt;|&lt;lib_b&gt; where lib_a precedes lib_b in alphabetical ordering.</div> <br> <div>Some of the examples are truncated for better readability.</div> <br> <div>GitHub source of the project: https://github.com/MeszarosGabor/SO_Post_Analyzer</div> <br> <div># Descriptions</div> <div>## `&lt;language&gt;`/all_`&lt;language&gt;`_so_posts.jsonl</div> <br> <div>JSONL file that contains the raw extracted Stack Overflow fields. Within a single JSON object:</div> <div>key: post_id,</div> <div>values:</div> <div>- post_type: 1 for question and 2 for answer</div> <div>- accepted_answer_id</div> <div>- date_posted</div> <div>- score</div> <div>- view_count</div> <div>- code_snippets</div> <div>- post_length</div> <div>- poster_id</div> <div>- last_actiivity</div> <div>- tags</div> <div>- number of comments</div> <div>- number of answers</div> <div>- parent id</div> <br> <div>Example:</div> <div>```</div> <div>{"72": ["1", "", "2008-08-01T13:38:27.133", "48", "2148", "&lt;p&gt;I want to format my existing comments as 'RDoc comments' so they can be viewed using &lt;code&gt;ri&lt;/code&gt;.&lt;/p&gt;\n\n&lt;p&gt;What are some recommended resources for starting out using RDoc?&lt;/p&gt;\n", "25", "2016-12-30T06:56:18.310", "&lt;ruby&gt;&lt;rdoc&gt;", "1", "2", ""]}</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_all_libs_dates.json</div> <br> <div>JSON file that lists the dates (with multiplicity, one for every post) when an individual library was mentioned in a post.</div> <br> <div>Example:</div> <div>```</div> <div>'FileUtils': ['2011-06-09',</div> <div>'2011-07-01',</div> <div>'2011-11-20',</div> <div>'2011-11-20',</div> <div>...</div> <div>'2013-09-04',</div> <div>'2020-05-08',</div> <div>'2021-02-25']</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_all_pairs_dates.json</div> <br> <div>JSON file that lists the dates (with multiplicity, one for every post) when a pair of libraries was mentioned in a post.</div> <br> <div>Example:</div> <div>```</div> <div>'mongo_mapper|sinatra': ['2010-09-12',</div> <div>'2011-12-30',</div> <div>'2012-02-23',</div> <div>'2012-09-04'],</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_libs_count.json</div> <br> <div>JSON file that lists the occurrence count of the individual libraries.</div> <br> <div>Example:</div> <div>```</div> <div>{</div> <div>'cairo': 4,</div> <div>'pango': 2,</div> <div>'radix': 1,</div> <div>}</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_pairs_count.json</div> <br><br> <div>JSON file that lists the co-occurrence count of the pairs of libraries.</div> <br> <div>Example:</div> <div>```</div> <div>'mongo_mapper|sinatra': 4,</div> <div>'fileutils|getoptlong': 1,</div> <div>'redis|rubygems': 24,</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_libs_first_dates.json</div> <br> <div>JSON file that lists the dates of the first appearances of individual libraries alongside the post id and poster id.</div> <br> <div>Example:</div> <div>```</div> <div>{</div> <div>'cairo': {'id': '6242589', 'poster_id': '784674', 'date': '2011-06-05'},</div> <div>}</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_pairs_first_dates.json</div> <div>JSON file that lists the dates of the first co-appearances of pairs libraries alongside the post id and poster id.</div> <br> <div>Example:</div> <div>```</div> <div>'rubygems|server': {'id': '3748309',</div> <div>'poster_id': '262808',</div> <div>'date': '2010-09-20'</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_`&lt;language&gt;`_code_count_list.json</div> <br> <div>JSON file that contains a single list of library counts in the posts (in chronological order) that contain *at least one* library import.</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_daily_post_stats.json</div> <div>JSON file that counts the number of posts on a given day, listed chronologically, containing dates *with at least one post*. Dictionary of key=date value=count(int) pairs.</div> <br> <div>Example:</div> <div>```{...</div> <div>'2011-09-03': 6,</div> <div>'2011-09-04': 3,</div> <div>'2011-09-05': 10,</div> <div>'2011-09-06': 5,</div> <div>'2011-09-07': 15,</div> <div>...}</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_`&lt;langugae&gt;`_post_stats.json</div> <br> <div>JSON file that lists the individual post metadata (sorted by post date).</div> <div>Fields:</div> <div>- post id,</div> <div>- post type,</div> <div>- list of imports</div> <div>- post date</div> <div>- poster id</div> <div>- score</div> <br> <div>Example:</div> <br> <div>```</div> <div>{'id': '1892176',</div> <div>'post_type': '1',</div> <div>'imports': ['mechanize', 'rubygems'],</div> <div>'date': '2009-12-12T03:31:43.823',</div> <div>'poster_id': '124685',</div> <div>'score': '5'},</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_time_based_new.jsonl</div> <br> <div>JSONL file that contains JSON objects (in chronological order) detailing post metadata.</div> <div>Fields:</div> <div>- post id,</div> <div>- post date</div> <div>- poster id (user id)</div> <div>- post type,</div> <div>- list of imports</div> <div>- list of novel libraries in post</div> <div>- list of novel pairs in post</div> <br> <div>Example:</div> <div>```</div> <div>{'post_id': '3543',</div> <div>'post_date': '2008-08-06T15:24:00.787',</div> <div>'user_id': '399',</div> <div>'post_type': '2',</div> <div>'imports': ['metric_fetcher', 'rake'],</div> <div>'new_libs': ['metric_fetcher', 'rake'],</div> <div>'new_pairs': ['metric_fetcher|rake']}</div> <div>```</div> <br> <div>## `&lt;language&gt;`/`&lt;language&gt;`_user_to_posts.json</div> <br> <div>JSON file that lists the post ids corresponding to a given user id. Keyed by user ids, values are list of post ids.</div> <br> <div>Example:</div> <div>```</div> <div>'303675': ['2941479'],</div> <div>'348325': ['2945141', '2956990', '2968924', '3832703'],</div> <div>'325477': ['2945228'],</div> <div>'27196': ['2949100', '3177217'],</div> <div>```</div> </div>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts".

<p>Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts". The data set consists of five folders: &ldquo;Codon_usage_amoebae_and_viruses&rdquo;, "Genome_annotations_amoebae", "Phylogenetic_trees_18S_amoebae", &ldquo;Phylogenomic_trees_amoebae&rdquo;, and "Viral_integration_detection_amoebae". The &ldquo;Codon_usage_amoebae_and_viruses&rdquo; folder contains for each amoeba host the calculated codon usage tables in the subfolder "codon_usage_table_host", the calculated codon usage preferences using different scores in the subfolder "codon_usage_scores_host", and the calculated codon usage preferences of giant viruses versus each host in the subfolder "codon_usage_scores_viruses_vs_host". The giant viruses in the subfolder "codon_usage_scores_viruses_vs_host" are organised by viral family and genus in separate sub-subfolders. The "Genome_annotations_amoebae" folder contains the generated genome annotations in different formats and the manually curated mitochondrial genome annotations for each amoeba host.&nbsp; The "Phylogenetic_trees_18S_amoebae" contains for the eukaryotic phyla <em>Discosea</em>, <em>Heterolobosea</em>, and <em>Tubulinea,&nbsp;</em>the 18S rRNA&nbsp;nucleotide alignments, distance matrices, and computed phylogenetic trees. The folder "Phylogenomic_trees_amoebae" contains for the eukaryotic clades <em>Amoebozoa</em> and <em>Discoba,&nbsp;</em>the protein alignment matrices and computed phylogenomic trees. The folder "Viral_integration_detection_amoebae" contains the MCP databases used (fasta file, alignment file, HMM profile and DIAMOND BLASTX database) and the MCP sequences detected in this study and the blast results of these.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Quantitative results of the analysis of human bioengineered tissues corresponding to the work "Development of novel squid gladius biomaterials for cornea tissue engineering"

<p>This dataset corresponds to the quantitative data generated in the work entitled "Development of novel squid gladius biomaterials for cornea tissue engineering".</p> <p>Cornea tissue engineering is strictly dependent on the development of biomaterials fulfilling the strict biocompatibility, biomechanical and optical requirements of this organ. In this work, we have generated novel biomaterials from the squid gladius (SG) and their application in cornea tissue engineering was evaluated. Results revealed that the native SG (N-SG) was biocompatible in laboratory animals, although a local inflammatory reaction was driven by the material. Cellularized biomaterials (C-SG) demonstrated that the SG provides an adequate substrate for cell attachment and growth, and corneal epithelial cells cultured on this biomaterial were able to express crystallin alpha, a marker for this type of cells. Biomechanical analyses showed that N-SG biomaterials have higher Young modulus and lower traction deformation than control native corneas (CTR), and C-SG showed similar Young modulus than CTR. Analysis of the optical properties of these samples revealed that the diffuse transmittance of N-SG and C-SG were higher than CTR, with the diffuse reflectance showing the opposite behavior. These results confirm the putative usefulness of this abundant marine-derived biomaterial that can be obtained as a byproduct of the fishing industry.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record