Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,199

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,199 results for “alignment”

Learn how ShareScore rates datasets ↗
zenodo44/100

MACSE Barcode Alignments

<p>This page provides alignments of barcoding sequences obtained in April 2020 by using the <a href="https://github.com/ranwez/MACSE_V2_PIPELINES/tree/master/MACSE_BARCODE">MACSE barcoding pipelines</a> on sequences collected from the BOLD database. The pipeline is fully described in:</p> <ul> <li>Fr&eacute;d&eacute;ric Delsuc and Vincent Ranwez (2020). Accurate alignment of (meta)barcoding data sets using MACSE. In Scornavacca, C., Delsuc, F., and Galtier, N., editors, Phylogenetics in the Genomic Era, chapter No. 2.3, pp. 2.3:1-2.3:31. No commercial publisher | Authors open access book. ( <a href="https://hal.inria.fr/PGE/hal-02541199">hal-02541199</a> ). (doi:<a href="https://doi.org/10.5281/zenodo.14185826">10.5281/zenodo.14185826</a>).</li> </ul> <p>Details of the individual files are provided in the file README.html.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Speech/text alignments for Italian end-to-end TTS

<p>Here are 146 chapters of several audiobooks from Librivox (33:22:19.176) read by 33 italian speakers:</p> <ul> <li>19 Females<br>#LC (Lisa Caputo): 18 chapters - 8643 utterances&gt; 07:51:25.991<br>#EG (Enrica Giampieretti): 16 chapters - 8473 utterances&gt; 07:04:47.044<br>#MT (Mariateresa): 7 chapters - 2735 utterances&gt; 02:57:0.487<br>#MR (Mariarosa): 2 chapters - 1150 utterances&gt; 00:55:59.825<br>#FA (Fabiola) 2 chapters - 324 utterances&gt; 00:13:36.635<br>#NI (Nicole Grassi) 2 chapters - 396 utterances&gt; 00:19:26.222<br>#SP (Simona Pagliari) 7 chapters - 557 utterances&gt; 00:26:58.920<br>#MM (Marzia Marianera) 2 chapters - 523 utterances&gt; 00:28:18.365<br>#ANGE (Angelina) 2 chapters - 246 utterances&gt; 00:15:45.223<br>#LAURA (Laura) 2 chapters - 899 utterances&gt; 00:38:33.592<br>#FG (Filippo Gioachin): 4 chapters - 1358 utterances&gt; 00:58:59.194<br>#GIOEMILY (?) 2 chapters - 743 utterances&gt; 00:33:26.872<br>#CAPI (Silvia di Simone) 1 chapters - 482 utterances&gt; 00:29:16.553<br>#ALLIE (Allie Cingi) 2 chapters - 838 utterances&gt; 00:50:31.221<br>#CAIMMA () 1 chapters - 328 utterances&gt; 00:19:26.483<br>#DOLCINEA () 1 chapters - 393 utterances&gt; 00:19:19.275<br>#MGT (Maria Grazia Tundo) 1 chapters - 290 utterances&gt; 00:10:41.887<br>#FR (Francesca Roma) 2 chapters - 246 utterances&gt; 00:16:16.611<br>#PETULA () 3 chapters - 251 utterances&gt; 00:16:27.894</li> <li>14 Males:<br>#RF (Riccardo Fasol): 3 chapters - 961 utterances&gt; 01:03:59.329<br>#RC (Roberto Confini): 3 chapters - 1238 utterances&gt; 00:42:2.761<br>#SB (Sergio Baldelli): 2 chapters - 910 utterances&gt; 00:54:57.145<br>#DA (Daniele) 2 chapters - 420 utterances&gt; 00:24:36.827<br>#RECL (Renzo Clerico) 3 chapters - 778 utterances&gt; 00:37:31.304<br>#STRALF (?) 3 chapters - 1392 utterances&gt; 00:51:23.847<br>#PAOLO (?) 2 chapters - 484 utterances&gt; 00:34:3.345<br>#PIER (?) 1 chapters - 320 utterances&gt; 00:14:10.860<br>#AB (Andrea Briglia) - 31 chapters - 863 utterances&gt; 00:40:26.526<br>#SIRJOE (Sergio Bersanetti) 2 chapters - 626 utterances&gt; 00:31:39.754<br>#KIUKKO (Luigi Chiaro) 1 chapters - 368 utterances&gt; 00:17:48.204<br>#AZ (Francesco Montana) 1 chapters - 287 utterances&gt; 00:14:39.814<br>#BM (Beniamino Massimo) 1 chapters - 415 utterances&gt; 00:16:6.842<br>#ML (Mirko Lamberti) 1 chapters - 252 utterances&gt; 00:12:27.594<br><br></li> <li>and a dictionnary of 14969 words with aligned phones</li> </ul> <p>Sources:</p> <ul> <li>Audios are from <a href="https://librivox.org/">Librivox</a></li> <li>Aligned original texts are from diverses sources including&nbsp;<a href="https://www.intratext.com">Intratext</a>, <a href="https://it.wikisource.org">wikisource</a>, <a href="https://www.pirandelloweb.com">pirandelloweb</a>, etc</li> </ul> <p>Each .wav file (sampled at 22050Hz) corresponds to one entire chapter. The format of the filenames is:<br>{author's acronym}_{book's acronym}_{reader's acronym}_{volume's number}_{chapter's number}</p> <p>The IT.csv file gives text (or sometimes, phones) and signal alignments for utterances in 4 fields separated by '|': {filename}|{start_ms}|{end_ms}|{text or phonetic content}. Most utterances are separated by at least a pause of 400ms (exceptionally less when phonation exceeds 11s). The intervals [start_ms:end_ms] comprise leading and trailing silences of 130ms (since wavs are entire chapters, these silences are "true" ambient silences).</p> <p>When phonetic alignment has been performed, 1 additional field has been added: {aligned phones}. Each input character or phone has a corresponding aligned phone. Note that all aligned utterances start and end with an aligned silence of 130ms. The set of aligned phones comprises:</p> <ul> <li>The set of input phones</li> <li>The silence: '__'</li> <li>The symbol '_' for silent characters, e.g. "occhi" is aligned with 'o^1 k: _ _ i'</li> </ul> <p>Text is in UTF8. '&laquo;&raquo;','&mdash;', '~','""','()','[]' are respectively used for speaking quotes, turn switches, three dots, quoted expression, aside quotes, notes. Because of rare occurrences, '&ouml;' has been transcribed as 'oe'. Paragraphs (two consecutive carriage returns in the original text) are cued by a special character '&sect;'. It usually ends an utterance but could be used within an utterance if its associated pause is too short.</p> <p>Part of text under clear emphasis is surrounded by "#"</p> <p>When available, phonetic content is given per word in curly brackets '{}'. We use 39 phonetic symbols:</p> <ul> <li><strong>oral vowels</strong>: a (f<strong>a</strong>), e (v<strong>e</strong>), e^ (<strong>e</strong>d), i (r<strong>iz</strong>), u (t<strong>u</strong>), o (un<strong>o</strong>), o^ (c<strong>o</strong>n)</li> <li><strong>loan vowels &amp; diphtongs: </strong>a&amp;i and x^ (t<strong>i</strong>m<strong>er</strong>),</li> <li><strong>semi-vowels</strong>: h (g<strong>h</strong>etto.), w (q<strong>u</strong>el), j (va<strong>j</strong>)</li> <li><strong>consonants</strong>: p (vespa), t (<strong>t</strong>u), k (<strong>c</strong>alde), b (<strong>b</strong>uon), d (<strong>d</strong>isse), g (<strong>g</strong>razie), f (<strong>f</strong>ame), s (<strong>s</strong>auna) , s^ (<strong>sc</strong>ia), v (<strong>v</strong>erde), z (ro<strong>s</strong>a), z^ (<strong>j</strong>udo), r (<strong>r</strong>izo), l (<strong>l</strong>etto), l^ (e<strong>gl</strong>i), m (<strong>m</strong>apo), n (<strong>n</strong>uda), n~ (pu<strong>gn</strong>i)</li> <li><strong>long/double consonants are suffxed by ":",</strong> e.g.&nbsp; p: (zu<strong>pp</strong>a)</li> <li><strong>primary stress</strong> if any is noted "1" and appended to the vowel, e.g. a1 g a p e (agape)</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Hepatocystis alignments and variant calls

<p>This contains BAM files of <em>Hepatocystis</em> reads identified in <em>Papio</em> and <em>Chlorocebus</em> samples mapped to the <em>Hepatocystis</em> reference genome as well as major and minor allele calls for cHEP and pHEP called with ANGSD, both individually and jointly. Nucleotide alignments have been added in the latest version.</p>

opencc-by-sa-4.0Jun 2024View details →
zenodo44/100

HOME-Alcar: Aligned and Annotated Cartularies

<p>The HOME-Alcar (Aligned and Annotated Cartularies) corpus was produced as part of the European research project HOME History of Medieval Europe (https://www.heritageresearch-hub.eu/project/home/), led under the coordination oflinebreakof Institut de Recherche et d&#39;Histoire des Textes (PI: D. Stutzmann), with the Universitat Politecnica de Valencia (PI: E. Vidal), the National Archives of the Czech Republic in Prague (PI: J. Kreckova), and Teklia SAS (PI: C. Kermorvant)<br> The HOME-Alcar (Aligned and Annotated Cartularies) corpus is a resource created to train Handwritten Text Recognition (HTR) and Named Entity Recognition (NER), and presents a collection of<br> (i) digital images of 17 medieval manuscripts;<br> (ii) scholarly editions thereof;<br> (iii) coordinates linking images and text at line level;<br> (iv) annotations of Named Entities (place and person names).<br> The 17 medieval manuscripts in this corpus are cartularies, i.e. books copying charters and legal acts, produced between the 12th and 14th centuries.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

KMA Mapping and alignment statistics : livestock fecal metagenomes against ResFinder and genomes

<p>Three zip archives are included used in the analysis of the European livestock resistome.</p> <p>Two of them contain &#39;mapstat&#39; files produced by the KMA software using the &#39;extended features&#39; flag.<br> Each mapstat file thus summarize the mapping and alignment statistics when using KMA on a metagenome against a database.</p> <p>The last archive contains the &#39;refdata&#39; file used to annotate the genomic mapstat hits. It encodes the taxonomic affilication of sequences hit by one or more samples.<br> &nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Aligned translation of Artemidorus Onir. book V

<p>Aligned translation of Artemidorus' Oneirocritica Book 5 in 95 chapters and a prologue divided into four sections. Part of the Open Projects in Digital Classics at the College of Letters and Sciences of the State University of S&atilde;o Paulo in Araraquara, S&atilde;o Paulo, Brazil. That is a second version of the translation. It was aligned on the Ugarit Platform and is visible at <a href="http://ugarit.ialigner.com/userProfile.php?userid=15&amp;tgid=9056">https://ugarit.ialigner.com/userProfile.php?userid=15&amp;tgid=14643</a>. The Greek text source was the digitized Pack's 1963 edition from CTS Perseids: urn:cts:greekLit:tlg0553.tlg001.1st1K-grc1:5. The Portuguese text is the revised translation (urn:cts.greekLit:tlg:0553.tlg001.ferreira2:5) of a previous&nbsp;<a href="https://www.culturaacademica.com.br/catalogo/oneirokritika-de-artemidoro-de-daldis-seculo-ii-d-c/">ebook</a>&nbsp;published by Cultura Acad&ecirc;mica in 2014 (digitized and available at https://furman-editions-in-progress.github.io/UNESP_FU/ as urn:cts:greekLit:tlg0553.tlg001.ferreira1:5).</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Alignment used in "A phylogenomically informed five-order system for the closest relatives of land plants"

<p>Alignment that served as the basis for the phylogenomic analyses presented in &quot;A phylogenomically informed five-order system for the closest relatives of land plants&quot; &mdash; preprint on bioRxiv&nbsp;doi: https://doi.org/10.1101/2022.07.06.499032</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Polymerase Structure Alignments - v4 Update (July 2022)

<p>v4&nbsp;update (July 2022) of superposed viral polymerase structures using the methodology described in&nbsp;<br> <em>A Comprehensive Superposition of Viral Polymerase Structures</em>&nbsp;by Olve Peersen, <a href="https://www.mdpi.com/1999-4915/11/8/745">Viruses 2019 11(8) 745</a>.</p> <p>&nbsp; &bull; &nbsp;<strong>_All_PDBs_v4.zip</strong> contains a single directory with coordinate files for 917 structures taken from 628 PDB entries. &nbsp;Uncompressed size is 1.6&nbsp;GB.</p> <p>&nbsp; &bull; <strong>Polymerases_v4.zip</strong> contains the complete set of files from the superpositions, including the _All_PDBs/ directory. &nbsp; Uncompressed size is 5.7&nbsp;GB.</p> <p>---------------------------------------------------------------------------------------------------------------<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; POLYMERASE STRUCTURES community page at Zenodo.org&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;v4 Release &nbsp;July 19, 2022</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 917 structures from 628 PDB entries</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; http://www.zenodo.org/communities/pols/&nbsp;<br> =========================================================================================</p> <p>This directory contains a large collection viral polymerase structure coordinates that&nbsp;have been superposed using the methods described in Peersen, Viruses 2019, 11(8), 745.&nbsp;</p> <p>They are released into the public domain under the Creative Commons Attribution (CC BY)&nbsp;license. &nbsp;</p> <p>Please cite the following two references for any derivative work:</p> <p>&nbsp; &nbsp;1) Reference to the publication describing the methodology:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Peersen, OB<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; A Comprehensive Superposition of Viral Polymerase Structures.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Viruses 2019, 11(8), 745 &nbsp;&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; https://doi.org/10.3390/v11080745</p> <p>&nbsp; &nbsp;2) DOI reference that always resolves to the latest release of the dataset:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Peersen, OB<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Polymerase Structure Alignments<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; https://doi.org/10.5281/zenodo.3361874</p> <p><br> ---------------------------------------------------------------------------------------------<br> Olve Peersen &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Olve.Peersen@ColoState.edu<br> Professor<br> Biochemistry &amp; Molecular Biology<br> 1870 Campus Delivery<br> Colorado State University<br> Ft. Collins, CO &nbsp;80523-1870<br> -------------------------------------------------------------------------</p> <p>=============================================================================</p> <p>v4&nbsp;&nbsp; &nbsp;220719 - Many updates and several news sets, as outlined below.&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp; The total count is now 917 polymerase structures from 628 PDB entries</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;2-rddp: Foamyvirus FOAM added, but rest of set left unchanged</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;3-dsdn: Simplified as hsv1 was pulled into its own set (see below)</p> <p><br> &nbsp;&nbsp; &nbsp;NEW&nbsp;&nbsp; &nbsp;alph: Alphaviruses - &nbsp;&nbsp; &nbsp; RRV: 7f0s,<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;SINV: 7vb4, 7vw5</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;aren:&nbsp;&nbsp; &nbsp;Lassa:&nbsp;&nbsp; &nbsp;7ckl, 7ela, 7och, 7oe3, 7oe7, 7oea, 7oeb,<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;7ojj, 7ojk, 70jl, 7ojn<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Machupo: 7ckm, 7el9, 7elb, 7elc, 7vgq, 7vh1, 7vh2, 7vh3</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;buny:&nbsp;&nbsp; &nbsp;RVF: &nbsp; 7eei<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;LaCrosse: 7ori, 7orj, 7ork, 7orl, 7orm, 7orn, 7oro</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;coro: &nbsp;&nbsp; &nbsp;7dfg, 7dfh, 7doi, 7dok, 7dte, 7ed5, 7egq, 7eiz, 7krn, 7kro, 7krp,&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;7oyg, 7ozu, 7ozv, 7rdx, 7rdy, 7rdz, 7re0, 7re1, 7re2, 7re3, 7thm&nbsp;</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;fluv:&nbsp;&nbsp; &nbsp;7nha, 7nhc, 7nhx, 7ni0, 7nik, 7nil, 7nir, 7nis, 7nj3, 7nj4, 7nj5,&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;7nj7, 7nk1, 7nk2, 7nk4, 7nk6, 7nk8, 7nka, 7nkc, 7nki, 7nkr, 7z42,&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;7z43, 7z43, 7z4o, 7z4o</p> <p>&nbsp;&nbsp; &nbsp;NEW&nbsp;&nbsp; &nbsp;foam: Foamyviruses - 7o0g, 7o0h, 7024<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Created FOAM.pdb from core of 7o0g.pdb, which was&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;superposed on rddp set, but the orientations of other&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;members of that set (HIV,MMLV,TERT) were not updated.</p> <p>&nbsp;&nbsp; &nbsp;NEW&nbsp;&nbsp; &nbsp;hsv1: Herpes Simplex Virus - 2gv9, 7luf</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pest: Classic Swine Fever Virus: 7ekj</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;poxv: 7amv, 7aof, 7aoh, 7aoz, 7ap8, 7ap9</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;rb69: 2atq, 7f4y</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;tavp: 7om2, 7om6, 7om7, 7om9, 7oma<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Data for "Hot-carrier transfer across a nanoparticle-molecule junction: The importance of orbital hybridization and level alignment"

<p>This upload includes the data presented and analyzed in the article &quot;Hot-carrier transfer across a nanoparticle-molecule junction: The importance of orbital hybridization and level alignment&quot; by Jakub Fojt, Tuomas P. Rossi, Mikael Kuisma, and Paul Erhart.</p> <p>The codes for reproducing the data are provided at <a href="https://doi.org/10.5281/zenodo.7118376">doi:10.5281/zenodo.7118376</a>.</p> <p>See <em>README.md</em> in <em>data.zip</em> for a detailed description.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Platynereis dumerilii - Aligned serial sections of the posterior segments

<p>Aligned serial semi-thin sections (1&micro;m) of the posterior most segments of <em>Platynereis dumerilii.</em></p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Elliptical Alignment Holes Enabling Accurate Direct Assembly of Microchips to Standard Waveguide Flanges at sub-THz Frequencies - Dataset

<p>Current waveguide flange standards do not allow for the accurate fitting of microchips, due to the large mechanical tolerances of the flange alignment pins and the brittle nature of Silicon, requiring greatly oversized alignment holes on the chip to fit worst-case fabrication tolerances, resulting in unacceptably large misalignment error for sub-THz frequencies. This paper presents, for the first time, a new method for directly aligning micromachined Silicon chips to standard, i.e. unmodified, waveguide flanges with alignment accuracy significantly better than the waveguide-flange fabrication tolerances, through the combination of a tightly-fitting circular and an elliptical alignment hole on the chip. A Monte Carlo analysis predicts the reduction of the mechanical assembly margin by a factor of 5.5 compared to conventional circular holes, reducing the potential chip misalignment from 46 μm to 8.5 μm for a probability of fitting of 99.5%. For experimental verification, micromachined waveguide chips using either conventional (oversized) circular or the proposed elliptical alignment holes were fabricated and measured. A reduction in the standard deviation of the reflection coefficient by a factor of up to 20 was experimentally observed from a total of 200 measurements with random chip placement, exceeding the<br> expectations from the Monte Carlo analysis. To our knowledge, this paper presents the first solution for highly accurate assembly<br> of micromachined waveguide chips to standard waveguide flanges, requiring no custom flanges or other tailor-made split blocks.<br>  </p>

opencc-by-nc-4.0Jun 2017View details →
zenodo44/100

Aligned DNA sequence matrix for phylogenetic analyses in the article "A new glassfrog of the genus Centrolene (Amphibia: Centrolenidae) from the Subandean Kutukú Cordillera, eastern Ecuador"

<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "A new glassfrog of the genus Centrolene (Amphibia: Centrolenidae) from the Subandean Kutuk&uacute; Cordillera, eastern Ecuador"</p> <p>The matrix is in NEXUS format and has 6626 bp and 239 terminals.</p> <p>Partitions are as follows:</p> <div> <div>charset 12S = 1-967;</div> <div>charset 16S = 968-2130;</div> <div>&nbsp;</div> <div>charset BNDFcodonPos1 = &nbsp;2133-2829\3;</div> <div>charset BNDFcodonPos2 = &nbsp;2131-2830\3;</div> <div>charset BNDFcodonPos3 = &nbsp;2132-2828\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset ND1codonPos1 = &nbsp;2832-3786\3;</div> <div>charset ND1codonPos2 = &nbsp;2833-3787\3;</div> <div>charset ND1codonPos3 = &nbsp;2831-3788\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset CXCR4codonPos1 = &nbsp;3790-4144\3;</div> <div>charset CXCR4codonPos2 = &nbsp;3791-4142\3;</div> <div>charset CXCR4codonPos3 = &nbsp;3789-4143\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset cmyccodonPos1 = &nbsp;4145-4547\3;</div> <div>charset cmyccodonPos2 = &nbsp;4146-4548\3;</div> <div>charset cmyccodonPos3 = &nbsp;4147-4549\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset POMCcodonPos1 = &nbsp;4551-5160\3;</div> <div>charset POMCcodonPos2 = &nbsp;4552-5161\3;</div> <div>charset POMCcodonPos3 = &nbsp;4550-5162\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset RAG1codonPos1 = &nbsp;5163-5616\3;</div> <div>charset RAG1codonPos2 = &nbsp;5164-5617\3;</div> <div>charset RAG1codonPos3 = &nbsp;5165-5618\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset SLC8A1codonPos1 = &nbsp;5620-6160\3;</div> <div>charset SLC8A1codonPos2 = &nbsp;5621-6158\3;</div> <div>charset SLC8A1codonPos3 = &nbsp;5619-6159\3;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>charset SLC8A3codonPos1 = &nbsp;6162-6627\3;</div> <div>charset SLC8A3codonPos2 = &nbsp;6163-6625\3;</div> <div>charset SLC8A3codonPos3 = &nbsp;6161-6626\3;</div> </div> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Alignments and ML trees of cassava brown streak virus and Ugandana cassava brown streak virus polyprotein nucleotide sequences

<p>Alignments of full and nearly full polyprotein-length nucleotide sequences from GenBank for the two ipomoviruses that cause cassava brown streak disease, in fasta format.&nbsp; Separate alignments for 67 cassava brown streak virus sequences and 81 Ugandan cassava brown streak virus sequences are provided, as well as a combined alignment of 148 sequences.&nbsp; Alignments were created with MUSCLE and then modified by eye in AliView.</p> <p>Also, two tree files (in nexus) format are supplied, resulting from a maximum likelihood analysis with IQTree on each of the two single-species datasets. Support for nodes with aLRT and 100 actual bootstrap replicates are provided (aLRT/bootstrap).</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Data for Weighted Manifold Alignment using Wave Kernel Signatures for Aligning Medical image Datasets

<p>Data used in MRI experiments in paper &#39;Weighted Manifold Alignment using Wave Kernel Signatures for Aligning Medical image Datasets&#39;. For each volunteer, breath-hold data (folder bhs) and dynamic free-breathing (folder dyn) data is provided in NIFTI format.</p>

opencc-by-4.0Feb 2019View details →
zenodo44/100

Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method

<p>This dataset contains a GitHub repository containing all the data, analysis, Nextflow workflows and Jupyter notebooks to replicate the manuscript&nbsp;titled &quot;Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method&quot;.</p> <p>It also contains the Multiple Sequence Alignments (MSAs) generated and well as the main figures and tables from the manuscript.</p> <p>The repository is also available at GitHub (https://github.com/cbcrg/dpa-analysis) release `v1.2`.</p> <p>For details on how to use the regressive alignment algorithm, see the T-Coffee software suite (https://github.com/cbcrg/tcoffee).</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Genome alignments for the project "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol" - Protocol optimization

<p>Genome alignments for data generated in the project &quot;<em>Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol &ndash; Optimization of the protocol.</em>&quot; Files names indicate unique identifiers of MOIRAI workflow runs, with the following structure: library name, dot, workflow ID (OP-WORKFLOW-CAGEscan-short-reads-v2.0.), dot, timestamp. The raw (FASTQ) data of each library is also deposited in Zenodo (<a href="https://doi.org/10.5281/zenodo.250156">10.5281/zenodo.250156</a>). Library names correspond to the following runs:</p> <ul> <li>&nbsp;NC33: 151007_M00528_0161_000000000-AEBDC</li> <li>&nbsp;NC37: 151204_M00528_0173_000000000-AEBEF</li> <li>&nbsp;NC38: 151211_M00528_0175_000000000-AE9PJ</li> <li>&nbsp;NC39: 160122_M00528_0185_000000000-AEB18</li> <li>&nbsp;NC42: 160302_M00528_0192_000000000-AELYK</li> </ul> <p>This data can be analysed using the &quot;CAGEr&quot; software package available from Bioconductor.&nbsp; The &quot;multiplex_files.zip&quot; file contains tables indicating which samples are biological replicates of each other or negative controls.</p>

opencc-zeroJul 2019View details →
zenodo44/100

MaSS - Multilingual corpus of Sentence-aligned Spoken utterances

<p><strong>Abstract</strong></p> <p>The CMU Wilderness Multilingual Speech Dataset is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for potentially 700 languages. However, the fact that the source content (the Bible), is the same for all the languages is not exploited to date. Therefore, this article proposes to add multilingual links between speech segments in different languages, and shares a large and clean dataset of 8,130 para-lel spoken utterances across 8 languages (56 language pairs).We name this corpus MaSS (Multilingual corpus of Sentence-aligned Spoken utterances). The covered languages (Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish) allow researches on speech-to-speech alignment as well as on translation for syntactically divergent language pairs. The quality of the final corpus is attested by human evaluation performed on a corpus subset (100 utterances, 8 language pairs).</p> <p><a href="https://arxiv.org/pdf/1907.12895.pdf">Paper </a>| <a href="https://github.com/getalp/mass-dataset">GitHub Repository</a>&nbsp;containing&nbsp;the scripts needed to build the data set from scratch (if needed)</p> <p><strong>Project structure</strong></p> <p>This repository contains 8 Numpy files, one for each featured language, pickled with Python 3.6. Each line corresponds to the spectrogram of the file mentioned in the file <em>verses.csv</em>. There is a direct mapping between the ID of the verse and its index in the list (thus verse with ID 5634 is located at index 5634 in the Numpy file). Verses not available for a given language (as stated by the value &quot;Not Available&quot; in the CSV file) are represented by empty lists in the Numpy files, thus ensuring a perfect verse-to-verse alignement between each file.</p> <p>Spectrogram were extracted using Librosa with the following parameters:</p> <pre><code>Pre-emphasis = 0.97 Sample rate = 16000 Window size = 0.025 Window stride = 0.01 Window type = 'hamming' Mel coefficients = 40 Min frequency = 20</code></pre> <p>&nbsp;</p>

openmit-licenseJul 2019View details →
zenodo44/100

Aligned Cox1 and Cob sequences from Oikopleura dioica and other tunicates.

<p>Supporting data for the manuscript &laquo; <em>Widespread use of the &ldquo;ascidian&rdquo; mitochondrial genetic code in tunicates</em> &raquo; containing a) assemblies of the cytochrome oxidase subunit 1 (Cox1) and cytochrome b (Cob) mitochondrial genes from ESTs downloaded from the Oikobase database and b) protein alignment of these sequences and other tunicate sequences to study which genetic code is used in tunicate mitochondria.</p>

opencc-zeroOct 2019View details →
zenodo44/100

Dataset for "Deflected Mantle Flow and Shearing-aligned Lithospheric Melt under the Strike-slip Dead Sea Rift"

<p>Dataset 1: All of the individual shear-wave splitting measurements in the Dead Sea rift, including 1855 A and B measurements and 1088 Null measurements</p> <p>Dataset 2: Three component seismic waveforms used for shear-wave splitting measurement, for PKS, SKKS, and SKS, respectively</p> <p>Dataset 3: Earthquake catalogue with magnitude Mb 2.6 or above in the Dead Sea rift. Downloaded from the International Seismological Centre (https://www.isc.ac.uk/)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Reference: International Seismological Centre (2024), On-line Bulletin, [Dataset] doi:10.31905/D808B830</p> <p>Dataset 4: Holocene Volcano List. Downloaded from Global Volcanism Program (https://volcano.si.edu/volcanolist_holocene.cfm)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Reference: Global Volcanism Program, 2024. Volcanoes of the World (v. 5.2.2; 22 Aug 2024). Distributed by Smithsonian Institution, compiled by Venzke, E. [Database] doi:10.5479/si.GVP.VOTW5-2024.5.2</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

SSP-aligned projected European water withdrawal/consumption at 5 arcminutes

<p><u><span>Release 0.9.1 &ndash; What is new?</span></u></p> <p><span><span>&middot;<span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span></span></span><span>Industrial water withdrawals were initially overestimated due to a problem with the input data, but they are now fixed.</span></p> <p><span><span>&middot;<span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span></span></span><span>Historical water withdrawal covering 1960-2020 added. Consistency between historical and projected water withdrawals is maintained.</span></p> <p><u><span>Contents and naming conventions</span></u></p> <p><span>Annual European water withdrawal: </span></p> <p><span>{scenario}_{sector} _year_millionm3_5min_Europe_{from_year}_{to_year}.nc</span></p> <p><span>Scenarios: historical, ssp1, ssp2, ssp3, ssp5</span></p> <p><span>Sector: dom, ind; stand for domestic/industrial</span></p> <p><span><span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span>From/to_year: 1960/2020, 2020/2100; historical/ssp projections.</span></p> <p><span>Each file holds two variables: {sector}ww and {sector}wc representing gross and net (i.e. consumptive use) water withdrawal. For exmaple, for the industrial sector, there are two variables: <em>indww </em>and <em>indwc.</em><u> </u>The fraction &lsquo;<em>1 &ndash; indwc/indww</em> &lsquo; represents the share of return flows. </span></p> <p>-------------------------</p> <p>The dataset provides annual water withdrawal and consumption estimates for Europe at a spatial resolution of 5 arcminutes, covering the periods 1960-2020 (historical) and 2020-2100 for four SSPs (1, 2, 3, and 5). Below, we outline the procedure used to downscale the population projections to a 5-arcminute resolution and describe the main equations applied to project water withdrawal and consumption under different SSPs.</p> <p>The development of the high-resolution (5 arcminute) projected water withdrawal and consumption for Europe follows the methodology outlined by Wada et al. (2011a, 2011b). This new release incorporates new projections for population, GDP per capita, and urbanization patterns from the latest SSP database (v3.0.1; available at <a href="https://data.ece.iiasa.ac.at/ssp/" target="_new">https://data.ece.iiasa.ac.at/ssp/</a>). Since this update is still in progress as of August 25<sup>th</sup>, 2024, some necessary input data are sourced from an earlier version of the SSP data (SSP 2013, see Table 1).&nbsp;&nbsp;<strong>All data and methods used to generate the results provided in this dataset are described in the Readme - Data and Methods file.</strong></p> <p><a name="_Ref175610392"></a>Table 1: Data availability in different versions of the SSP database as of August 25<sup>th</sup> 2024.</p> <table> <tbody> <tr> <td> <p><strong>Data</strong></p> </td> <td> <p><strong>SSP DatabaseVersion</strong></p> </td> <td> <p><strong>Module</strong></p> </td> </tr> <tr> <td> <p>Population</p> </td> <td> <p>SSP v3.0.1 2024</p> </td> <td> <p>Domestic</p> </td> </tr> <tr> <td> <p>GDP per capita</p> </td> <td> <p>SSP v3.0.1 2024</p> </td> <td> <p>Domestic/industrial</p> </td> </tr> <tr> <td> <p>Energy use per capita</p> </td> <td> <p>SSP 2013</p> </td> <td> <p>Industrial</p> </td> </tr> <tr> <td> <p>Electricity use per capita</p> </td> <td> <p>SSP 2013</p> </td> <td> <p>Industrial</p> </td> </tr> </tbody> </table>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record