Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,555
datasets available to search
ShareScore release 0.7.1
Dataset results
2,555 results for “catalogs”
The data for SKYSURF-5: Probing the Integrated Galaxy Light with a SDSS-SKYSURF Cross-Matched Catalog
<p>The SKYSURF Project (Windhorst et al. 2022) analyzes the extragalactic background light (both directly using sky background measurements and indirectly using galaxy counts) using the HST Archive. While HST images probe faint galaxies unseen by ground-based imaging, its small field of view prevents it from probing the large-scale structure around its observations.</p> <p>To supplement SKYSURF analysis, we cross-match SKYSURF pointings with SDSS observations able to probe the surrounding large-scale environment (Bhatia et al. 2024). The tables in this database include galaxies brighter than r=22.5 AB mag photometrically identified in SDSS, within +/-5 arcmin around a SKYSURF pointing.</p> <p>The tables in the Object_AB directory include information on all SDSS objects, organized by the HST camera and filter of the central pointing.</p> <p>The tables in the IGL directory include the total galaxy counts and integrated galaxy light (down to AB mag=22.5) for all SDSS objects surrounding a given SKYSURF image.</p>
THE HIGH CADENCE TRANSIENT SURVEY (HITS): Source, light-curve and classification catalogs
<p>The High Cadence Transient Survey (HiTS) aims to discover and study transient objects with characteristic timescales between hours and days, such as pulsating, eclipsing and exploding stars. This survey represents a unique laboratory to explore large <em>etendue</em> observations from cadences of about 0.1 days and to test new computational tools for the analysis of large data. This work follows a fully <em>Data Science</em> approach: from the raw data to the analysis and classification of variable sources. We compile a catalog of ~15 million object detections and a catalog of ~2.5 million light-curves classified by variability. The typical depth of the survey is 24.2, 24.3, 24.1 and 23.8 in <em>u</em>, <em>g</em>, <em> </em>r, and <em>i</em> bands, respectively. We classified all point-like non-moving sources by first extracting features from their light--curves and then applying a Random Forest classifier. For the classification, we used a training set constructed using a combination of cross-matched catalogs, visual inspection, transfer/active learning, and data augmentation. The classification model consists of several Random Forest classifiers organized in a hierarchical scheme. The classifier accuracy estimated on a test set is approximately 97%. In the unlabeled data, 3,485 sources were classified as variables, of which 1,321 were classified as periodic. Among the periodic classes we discovered with high confidence, 1 <span class="math-tex">\(\delta\)</span> scuti, 39 eclipsing binaries, 48 rotational variables, and 90 RR-Lyrae. For the non-periodic classes we discovered 1 cataclysmic variables, 630 QSO, and 1 supernova candidate.</p>
Catalog of laboratory parameters for the Lyman and Werner bands of molecular hydrogen
<p>Catalog of laboratory parameters for the Lyman and Werner bands of molecular hydrogen based on the compilation of Malec et al. (2010). The wavelength values are partly updated to the measurements from Table 11 and 12 in Bailly et al. (2010). The values for the oscillator strength f and the damping constant gamma are re-calculated following (Morton 2003).</p> <p> </p>
Extract from the Berlin State Library's Main Catalog
<p>The data set is based on the main catalog of the library. Currently, the following fields are extracted:</p> <ul> <li>title</li> <li>author (+ optional GND ID)</li> <li>publisher</li> <li>place of publication</li> <li>country of publication</li> <li>year of publication</li> </ul> <p>The extract has been created by the processPicaPlus script available <a href="https://github.com/elektrobohemian/StabiHacks">here</a>. Attention, some special characters might not have been extracted correctly in versions <1.0.0.</p> <p><em>Change Log:</em></p> <p>0.2.0 </p> <p>fixes various encoding issues for non-ASCII characters</p> <p>0.3.0 </p> <p>added year of publication; added separate files for languages: rus, pol, rum, cze, slo, gre with minor encoding issues</p> <p> </p> <p><em>Dataset Characteristics</em></p> <p>The following languages are available in separate data files:</p> <ul> <li>eng</li> <li>ger</li> <li>lat</li> <li>fre</li> <li>ita</li> <li>spa</li> <li>por</li> <li>dut</li> <li>swe</li> <li>dan</li> <li>nor</li> <li>ice</li> <li>fry</li> <li>rus*</li> <li>pol*</li> <li>rum*</li> <li>cze*</li> <li>slo*</li> <li>gre*</li> </ul> <p>*: Language file might be subject to character encoding issues.</p> <p>The other languages are present in the data set but have not been separated, i.e., they are combined in one data file:</p> <p>'fre', 'rus', 'pol', 'ger', 'eng', 'lit', 'dan', 'dut', 'spa', 'swe', 'ita', 'lat', 'nor', 'ind', 'bul', 'grc', 'fry', 'rum', 'cze', 'slo', 'bel', 'ice', 'fin', 'gre', 'hun', 'tur', 'enm', 'hrv', 'est', 'srp', 'roh', 'syr', 'wen', 'mal', 'afr', 'slv', 'mac', 'smi', 'nds', 'qmw', 'pra', 'oci', 'bre', 'san', 'alb', 'baq', 'non', 'ara', 'chm', 'per', 'cat', 'gmh', 'sla', 'arm', 'ukr', 'por', 'chu', 'heb', 'arc', 'gle', 'tib', 'lav', 'geo', 'crp', 'hin', 'mul', 'chi', 'epo', 'kor', 'kan', 'vot', 'csb', 'glg', 'kaz', 'frm', 'jpn', 'bur', 'srd', 'sal', 'ira', 'bos', 'mol', 'rom', 'tat', 'aze', 'yid', 'mar', 'mak', 'pli', 'rys', 'tgk', 'map', 'vie', 'tuk', 'oss', 'ota', 'tut', 'ben', 'sun', 'tir', 'bak', 'chv', 'ber', 'khm', 'may', 'pan', 'uzb', 'swa', 'kir', 'egy', 'dum', 'nep', 'cop', 'mon', 'tam', 'urd', 'zxx', 'wel', 'mis', 'ng', 'goh', 'dt', 'en', 'fao', 'fro', 'pus', 'kur', 'cus', 'hau', 'uig', 'sit', 'dt.', 'cpf', 'tgl', 'qoj', 'tag', 'raj', 'fiu', 'xal', 'kbd', 'udm', 'scr', 'gag', 'kas', 'scc', 'pro', 'tha', 'dar', 'dr', 'sna', 'ewe', 'de', 'dra', 'ang', 'ine', 'zza', 'und', 'ave', 'amh', 'crh', 'jav', 'cpe', 'akk', 'dsb', 'qce', 'guj', 'ltz', 'got', 'bua', 'peo', 'mdr', 'nob', 'ava', 'che', 'sux', 'kok', 'zap', 'nl', 'inc', 'sah', 'gem', 'law', 'bem', 'sin', 'qdo', 'hsb', 'som', 'lao', 'kam', 'kom', 'abk', 'roa', 'cau', 'ady', 'bat', 'mlt', 'sai', 'xho', 'paa', 'sot', 'bnt', 'lug', 'myn', 'kar', 'qhe', 'kin', 'zul', 'tsn', 'apa', 'nso', 'yao', 'yor', 'bih', 'nog', 'nap', 'loz', 'nbl', 'kon', 'nya', 'snh', 'chn', 'run', 'suk', 'fur', 'osa', 'bra', 'den', 'kpe', 'kal', 'tig', 'wol', 'gla', 'lad', 'mos', 'cre', 'krc', 'ge', 'fr', 'dak', 'fij', 'mad', 'srr', 'kum', 'her', 'nai', 'cel', 'inh', 'kro', 'hit', 'pal', 'tmh', 'tsw', 'bam', 'kab', 'kik', 'kua', 'lub', 'luo', 'nub', 'tem', 'znd', 'mai', 'tai', 'qkr', 'ful', 'man', 'lol', 'sag', 'tog', 'hai', 'arg', 'fat', 'nav', 'niu', 'ibo', 'ido', 'men', 'qju', 'gaa', 'vol', 'nah', 'mlg', 'nic', 'ijo', 'sus', 'orm', 'smo', 'mag', 'tyv', 'mnc', 'cos', 'mdf', 'kaa', 'dua', 'gez', 'ton', 'ven', 'snd', 'syc', 'nym', 'nia', 'sem', 'chg', 'fan', 'twi', 'mas', 'ina', 'ile', 'art', 'ori', 'qai', 'arw', 'mao', 'bas', 'kmb', 'tiv', 'bal', 'tar', 'tpi', 'abs', 'asm', 'qqa', 'iku', 'min', 'rup', 'tel', 'or', 'tah', 'aka', 'day', 'qqg', 'lah', 'lus', 'sio', 'oto', 'alg', 'shn', 'ndo', 'haw', 'tso', 'mus', 'cai', 'qev', 'new', 'zha', 'grn', 'khi', 'ssw', 'nde', 'bla', 'grb', 'mun', 'din', 'sam', 'mwr', 'cor', 'sat', 'cho', 'ger,', 'que', 'btk', 'glv', 'rar', 'jk', 'nno', 'cmc', 'mga', 'jw', 'iro', 'sog', 'hat', 'dzo', 'mkh', 'bik', 'ban', 'ilo', 'pam', 'ts', 'sme', 'myv', 'qnn', 'jpr', 'qte', 'yap', 'bis', 'sga', 'qkj', 'pap', 'ath', 'ipk', 'phi', 'sco', 'del', 'moh', 'iri', 'gae', 'ryl', 'our', 't--', 'grk', 'ssa', 'awa', 'efi', 'jrb', 'enk', 'kru', 'oji', 'arn', 'car', 'gsw', 'lez', 'war', 'ace', 'qrn', 'wln', 'ceb', 'aar', 'bug', 'kaw', 'chr', 'cpp', 'tet', 'aym', 'ces', 'hmo'</p>
Title, Author, Publisher, Place of Publication, and Language-related Network Graphs of the Berlin State Library Main Catalog
<p>The dataset contains graphs in GML, GraphML, and a simple JSON format.</p> <p>For each of the following languages:</p> <ol> <li>cze</li> <li>dan</li> <li>dut</li> <li>eng</li> <li>fre</li> <li>fry</li> <li>ger</li> <li>gre</li> <li>ice</li> <li>ita</li> <li>lat</li> <li>nor</li> <li>pol</li> <li>por</li> <li>rum</li> <li>rus</li> <li>slo</li> <li>spa</li> <li>swe</li> </ol> <p>two graphs are made available linking</p> <ul> <li>author, publisher, and place of publication</li> <li>author, publisher, place of publication, and title</li> </ul> <p>Additionaly, a third graph links authors and publishers to the language of publication (incl. year of the publication).</p> <p>The core statistics of each graph are outlined in <em>social_analysis_statistics.csv</em>. The smallest graph (fry, author_publisher_location) has 298 nodes and 264 edges, while the largest (ger, author_publisher_location_title) has 2,499,943 nodes and 3,950,900 edges.</p> <p>The language graphs spans all languages and has 1,706,273 nodes and 1,827,759 edges.</p> <p>All graphs have been created by a Python script available <a href="https://github.com/elektrobohemian/CulturalAnalytics/blob/master/SocialAnalysisStabikat.ipynb">here.</a></p>
Catalog Data for Prior-Informed AGN-Host Spectral Decomposition Using PyQSOFit
<p>This catalog contains 76,565 AGN-host decomposed spectral measurements for all quasars with z<0.8 in SDSS DR16Q. Our prior-informed decomposition method significantly improved the decomposition success rate from less than 60% to 94%. For the first time, we perform the AGN-host spectral decomposition on survey scale catalog.</p> <p>Our spectral decomposition results are highly consistent to those of HSC image decomposition. Our catalog suggests that an average host galaxy contribution at 5100A is 38.8%, which would lead to an overestimation of 0.215 dex in L5100 and 0.219 dex in black hole mass if the host is not removed. The Dn4000 and stellar velocity dispersion measurements from the decomposed host galaxy spectra are also provided.</p> <p>Please read this paper for more techinique details: <a href="https://arxiv.org/abs/2406.17598">arXiv: 2406.17598</a></p>
HELP Study Data Dictionary / Catalog of Items
<p>The <a href="https://doi.org/10.1136/bmjopen-2019-033391">HELP study</a> was a clinical trial in the form of a multicenter interventional randomized controlled trial (RCT) at five German university hospitals, conducted from 2020 to 2022. It aimed to enhance the clinical management of Staphylococcus bacteremia and used data both from Electronic Health Records (EHR), provided by German university hospital's Data Integration Centers in the HL7 FHIR format (German profiles of the Medical Informatics Initiative Core Data Set – <a href="https://www.medizininformatik-initiative.de/en/medical-informatics-initiatives-core-data-set">MII CDS</a>), as well as data from Electronic Case Report Forms (eCRF) used in the study for data which was too unstructered or not availabe in the EHR (hybrid data collection approach).<br>This dataset is a tabular listing and description of the data items used in the study and <em>serves (primarily) as a template for information and data modeling.</em></p>
Catalog for The MDW Hα Sky Survey: Data Release 0
<p>Here, we upload the source catalog for Data Release 0 (DR0) of the MDW Hα Sky Survey. This catalog of ~1.9 million sources is catalog-matched to the Pan-STARRS DR1 catalog, with a subset of 160k sources matched to the IGAPS survey of the Galactic plane. </p> <p>This catalog is related to the AJ manuscript titled "The MDW Hα Sky Survey: Data Release 0", currently undergoing review (submission ID: AAS55752R1).</p> <p>More data from DR0 (e.g. images, QA) can be accessed at <a href="https://mdw.astro.columbia.edu" target="_blank" rel="noopener">https://mdw.astro.columbia.edu</a>. Any use of DR0 data (including the source catalog) should include the following acknowledgement:</p> <blockquote> <p>Funding for the MDW Survey Project has been provided by the Michele and David Mittelman Family Foundation. David R. Mittelman, Dennis di Cicco, and Sean Walker are founding members of the survey and made possible the acquisition and reduction of the data. Columbia University Astronomy Department is responsible for the final data reduction, calibration, and dissemination of the survey data. All commercial rights for the use of the data are reserved.</p> </blockquote>
The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts
<p>This data continues with the development of the NPEGC Trinity <em>de novo</em> metatranscriptome assemblies from the protein data repository of <a href="../doi/10.5281/zenodo.10472589">The North Pacific Eukaryotic Gene Catalog</a>. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:<br><br><em>NPac.G1PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G2PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G3PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G3PA_diel.bf100.id99.nt.fasta.gz</em><br><em>NPac.D1PA.bf100.id99.nt.fasta.gz</em><br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.<br><br>These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: <a href="../records/7332796">The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3</a></p> <p>Key processing steps are sampled below with links to the detailed code on the main github code repository: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog">https://github.com/armbrustlab/NPac_euk_gene_catalog</a></p> <p><br>Code used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/nt_data/NPEGC.nt_kallisto_counts.sh">NPEGC.nt_kallisto_counts.sh</a><br><br>There are two main steps:<br>1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts<br>2. Map the short reads from environmental samples back to the assembly index</p> <p>As generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.</p> <p>The code in this template script was used for each project: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/nt_data/aggregate_kallisto_counts.R">aggregate_kallisto_counts.R</a><br>The output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: </p> <p><em>G1PA.raw.est_counts.csv.gz</em><br><em>G2PA.raw.est_counts.csv.gz</em><br><em>G3PA.raw.est_counts.csv.gz</em><br><em>G3PA_diel.raw.est_counts.csv.gz</em><br><em>D1PA.raw.est_counts.csv.gz</em></p>
PEI_catalog
<p>Catalog of peak event intensity (PEI) of short heavy precipitation events. Based on RADKLIM YW.V2017.002 binary data (Winterrath et al. 2018). The dataset contains precipitation duration levels from 5 minutes to 180 minutes. A shapefile of the RADKLIM grid is stored for easy localization of the event zone (radklim_grid.zip). </p> <p>Column explanation: (1) <em>timestamp</em> = event start time [YYYY-MM-DD hh:mm:ss], (2) <em>intensity_0</em> = precipitation intensity in the event center [mm/h], (3) <em>height_0</em> = precipitation height in the event center [mm], (4) <em>height_1</em> = mean precipitation height over first order neigbors (9 pixel) [mm], (5) <em>height_2</em> = mean precipitation height over second order neigbors (27 pixel) [mm], (6) <em>height_3</em> = mean precipitation height over third order neigbors (45 pixel) [mm], (7)<em> duration</em> = event duration level [min], (8) <em>extremity_ratio</em> = ratio of <em>height_0</em> and the threshold value of the corresponding <em>duration</em>, (9) <em>radklim_id</em> = RADKLIM grid ID (start in the upper left corner), (10) <em>radklim_y</em> = RADKLIM grid y [row], (11) <em>radklim_x</em> = RADKLIM grid x [col], (12) <em>utm32_y</em> = utm 32 y-center coordinate (EPSG:25832), (13) <em>utm32_x </em>= utm 32 x-center coordinate (EPSG:25832)</p>
GWTC-2.1: Deep extended-catalog of Compact Binary Coalescences Observed by LIGO and Virgo During the First Half of the Third Observing Run - Sensitivity of search pipelines to simulated signals
<p>Results of search pipelines (GSTLAL, MBTA, PYCBC, PYCBC BBH) used to identify candidates in <a href="https://dcc.ligo.org/LIGO-P2100063/public">GWTC-2.1</a> on a set of simulated signals corresponding to binary neutron star (BNS), neutron star black holes (NSBH), and binary black holes (BBH) signals. Additionally, we include a README file which provides information on how to read these files.</p>
HTRCatalogs: Dataset for historical catalogs HTR and Segmentation
<p>This release contains 465 xml files, and their corresponding images from a large corpus of 19th, 20th and 21th exhibition catalogs, manuscripts'fair catalogs and directories. The new catalogs added here were created using the HTR and segmentation models accessible in the repository. It includes a csv file describing the xml files and various tools to create a training dataset: differents bash scripts, a python programm to divide the xml files into testing, training and evaluation dataset and several fixed tests. A xsl transformation sheet is also accessible to delete the Entry and EntryEnd zones from the xml files in order to have a SegmOnto-like dataset. The xml files has been corrected since the 4.0 release thanks to the addition of a github action (SegmOntoKraken).</p>
Catalog of Repeating Earthquakes for Northern California, 1984-2014
<p>This catalog includes 27,675 repeating earthquakes grouped in 7,713 sequences in northern California for the years 1984-2014. The repeating earthquakes were identified by a comprehensive analysis of waveform similarity, relative event location, and relative size of all earthquakes recorded by the Northern California Seismic Network (NCSN). Details can be found in:</p> <p>Waldhauser, F., & Schaff, D. P. (2021). A comprehensive search for repeating earthquakes in northern California: Implications for fault creep, slip rates, slip partitioning, and transient stress. Journal of Geophysical Research: Solid Earth, 126, e2021JB022495. https://doi.org/10.1029/2021JB022495</p> <p> </p>
penguinlibros_esp_catalog
<p>Dataset created by webscrapping penguinlibros website for a delivery of the subject "tipología y ciclo de vida de los datos" of the "universidad oberta de catalunya".</p> <p>Contains information about a list of books listed in the website.</p>
HETDEX Public Source Catalog 1: 220 K Sources including over 50K Lyman Alpha Emitters From an Untargeted Wide-area Spectroscopic Survey
<p>We present the first publicly released catalog of sources obtained from the Hobby-Eberly Telescope Dark Energy Experiment (HETDEX). HETDEX is an integral field spectroscopic survey designed to measure the Hubble expansion parameter and angular diameter distance at 1.88 < z < 3.52 by using the spatial distribution of more than a million Lyα-emitting galaxies over a total target area of 540 deg2. The catalog comes from contiguous fiber spectra coverage of 25 deg2 of sky from January 2017 through June 2020, where object detection is performed through two complementary detection methods: one designed to search for line emission and the other a search for continuum emission. The HETDEX public release catalog is dominated by emission-line galaxies and includes 51,863 Lyα-emitting galaxy (LAE) identifications and 123,891 [O II]-emitting galaxies at z < 0.5. Also included in the catalog are 37,916 stars, 5,274 low-redshift (z < 0.5) galaxies without emission lines, and 4,976 active galactic nuclei. The catalog provides sky coordinates, redshifts, line identifications, classification information, line fluxes, [O II] and Lyα line luminosities where applicable, and spectra for all identified sources processed by the HETDEX detection pipeline. Extensive testing demonstrates that HETDEX redshifts agree to within ∆z < 0.02, 96.1% of the time to those in external spectroscopic catalogs. We measure the photometric counterpart fraction in deep ancillary Hyper Suprime-Cam imaging and find that only 55.5% of the LAE sample has an r-band continuum counterpart down to a limiting magnitude of r ∼ 26.2 mag (AB) indicating that an LAE search of similar sensitivity with photometric pre-selection would miss nearly half of the HETDEX LAE catalog sample.<br> <br> Two catalogs make up HETDEX Source Catalog 1:<br> <br> 1. The Source Observation Table: hetdex_sc1_vX.dat/.fits/.ecsv<br> With SPECTRA arrays included: hetdex_sc1_spec_vX.fits<br> <br> One row per source observation. The table provides basic coordinates/redshift/source information for each observation of a unique astornomical source. The larger file hetdex_sc1_spec_vX.fits contains the same info from the first table plus addition data units of spectral array data.<br> <br> 2. The Detection Information Table: hetdex_sc1_detinfo_vX.fits<br> One row per line or continuum detection. Bright sources can be comprised of multiple line or continuum emission. This catalog provides specific detection information such as line parameter info (S/N, line flux, line width)<br> <br> 3. Jupyter Notebook with access example: HETDEX_source_catalog_1.ipynb/.pdf/.html</p> <p>We request that the following acknowledgement be included in any paper using data or software from HETDEX public data releases:</p> <blockquote> <p>HETDEX is led by the University of Texas at Austin McDonald Observatory and Department of Astronomy with participation from the Ludwig-Maximilians-Universität München, Max-Planck-Institut für Extraterrestrische Physik (MPE), Leibniz-Institut für Astrophysik Potsdam (AIP), Texas A&M University, Pennsylvania State University, Institut für Astrophysik Göttingen, The University of Oxford, Max-Planck-Institut für Astrophysik (MPA), The University of Tokyo and Missouri University of Science and Technology.</p> <p>Observations for HETDEX were obtained with the Hobby-Eberly Telescope (HET), which is a joint project of the University of Texas at Austin, the Pennsylvania State University, Ludwig-Maximilians-Universität München, and Georg-August-Universität Göttingen. The HET is named in honor of its principal benefactors, William P. Hobby and Robert E. Eberly. The Visible Integral-field Replicable Unit Spectrograph (VIRUS) was used for HETDEX observations. VIRUS is a joint project of the University of Texas at Austin, Leibniz-Institut für Astrophysik Potsdam (AIP), Texas A&M University, Max-Planck-Institut fürExtraterrestrische Physik (MPE), Ludwig-Maximilians-Universität München, Pennsylvania State University, Institut für Astrophysik Göttingen, University of Oxford, and the Max-Planck-Institut fur Astrophysik (MPA).</p> <p>Funding for HETDEX has been provided by the partner institutions, the National Science Foundation, the State of Texas, the US Air Force, and by generous support from private individuals and foundations.</p> </blockquote> <p> </p>
Catalog of the Miyaoka Collection
<p>A catalog of <em>the Miyaoka Bunko</em>, a large-scale collection of research books with high rarity value mainly on Eskimo and North American languages, donated to the Language Research Institute of Tokyo University of Foreign Studies by world-famous linguist Professor Osahito Miyaoka.</p>
Global Biodiversity Information Facility (GBIF): an exhaustive list of gbif record ids, dataset keys, and their associated Occurrence IDs, Institution Code, Collection Codes and Catalog Numbers. hash://sha256/ea88f03a7bfd1ba853fdbea3203d54ab81ac3cdc8e8da7c96bbbba9c4b05d933 hash://md5/c49fe34785354847b37ea4509261e130
<p>The Global Biodiversity Information Facility (GBIF) indexes thousands of biodiversity datasets from Natural History Collections, citizen science initiatives (e.g., iNaturalist, eBird), and other sources. As part of the index process, GBIF associates at least two identifiers with indexed records: a record id (aka gbifID) and a dataset id (aka dataset key). These ids are central to do lookup, reference data, and package interpreted data products.</p> <p>This publication contains an exhaustive list of GBIF IDs and ids associated by their data providers as derived from:</p> <p>GBIF.org (01 March 2023) GBIF Occurrence Download https://doi.org/10.15468/dl.pk3trq</p> <p>The resource (size: ~260GB) provided by GBIF had content id hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 and was used to generate the resource included in this publication using</p> <pre><code class="language-bash">preston cat 'zip:hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97!/0015281-230224095556074.csv'\ | cut -f 1,2,3,37,38,39\ | gzip\ > gbifid.tsv.gz </code></pre> <p>with the content id of gbifid.tsv.gz (size: ~35GB) being hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8 .</p> <p>the first 10 lines of gbifid.tsv.gz as extracted via</p> <pre><code>preston cat --remote https://zenodo.org/record/7789866/files,https://linker.bio hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8\ | gunzip\ | head</code></pre> <p>are:</p> <pre><code>gbifID datasetKey occurrenceID institutionCode collectionCode catalogNumber 2997162320 c71c8000-9fc7-422c-804a-ce6abe751771 3399442 CEPEC CEPEC CEPEC00109669 2997162309 c71c8000-9fc7-422c-804a-ce6abe751771 2733085 CEPEC CEPEC CEPEC00000818 2997162317 c71c8000-9fc7-422c-804a-ce6abe751771 2733086 CEPEC CEPEC CEPEC00000888 2997162313 c71c8000-9fc7-422c-804a-ce6abe751771 3399443 CEPEC CEPEC CEPEC00109744 2997162306 c71c8000-9fc7-422c-804a-ce6abe751771 2733087 CEPEC CEPEC CEPEC00000889 2997162316 c71c8000-9fc7-422c-804a-ce6abe751771 3399440 CEPEC CEPEC CEPEC00109605 2997162324 c71c8000-9fc7-422c-804a-ce6abe751771 2733088 CEPEC CEPEC CEPEC00000890 2997162308 c71c8000-9fc7-422c-804a-ce6abe751771 3399441 CEPEC CEPEC CEPEC00109615 2997162303 c71c8000-9fc7-422c-804a-ce6abe751771 2733089 CEPEC CEPEC CEPEC00000891</code></pre> <p>Note that at time of writing, the html resource associated with the occurrence id 2997162320, and data set key c71c8000-9fc7-422c-804a-ce6abe751771 (extracted from of the first data row example above) are available via:</p> <p>https://gbif.org/occurrence/2997162320</p> <p>and</p> <p>https://gbif.org/dataset/c71c8000-9fc7-422c-804a-ce6abe751771</p> <p>respectively.</p> <p>This resource was initially created to help integrate with Bionomia (https://bionomia.net) to help associate people identifiers provided by bionomia to their original records via their GBIF ids. Bionomia re-uses GBIF records ids as a way to define links between records and the people (e.g., curators, collectors, identifiers) that worked on them. </p> <p>In other words, this resource provides a versioned translation table from the GBIF data universe (as defined by GBIF record ids, and dataset keys) to the data collections that exist (and evolve) independent of it. </p> <p>Note that the resource identified by hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 was not included in this publication it was too big (260GB) to fit. You may be able to retrieve the resource from its original location at https://api.gbif.org/v1/occurrence/download/request/0015281-230224095556074.zip .</p>
A catalog of associated, machine-learning-derived phase arrival times for ten days of seismic data in the Yellowstone region
<p>This dataset contains the associated phase picks and event information from applying a deep learning phase picker to continuous data recorded over March 25 – April 3, 2014, on 20 three-component stations and 14 vertical-component stations in the Yellowstone region. This 10-day period contains an M<sub>w</sub> 4.8 event, the largest earthquake in the Yellowstone region since 1980. The catalog and deep learning phase picker are described in Armstrong et al. (submitted).</p> <p>The arrivals were associated using the method described by Baker et al. (2021) and located using HypoInverse2000 (Klein, 2002). There are 1,053 events in this catalog, including 855 that were previously unidentified. Events that also appear in the University of Utah Seismograph Stations catalog have an event identifier (evid) beginning with “6”, while new events begin with “9”. </p> <p>Columns include:</p> <ul> <li>A simple event number</li> <li>the network, station, channel, and location code for the arrival time</li> <li>the arrival time in UTC (arrival_time) and Unix (arrival_time_epoch) format</li> <li>any static correction applied to the arrival time</li> <li>the P-pick first motion polarity as determined by a machine learning model - up (1), down (-1), or unknown (0)</li> <li>the arrival time residual </li> <li>the take off angle in degrees </li> <li>the event latitude and longitude in degrees</li> <li>the event depth in km</li> <li>the event origin time in UTC (origin_time) and Unix (origin_time_epoch) format</li> <li>the azimuthal gap of the event in degrees</li> <li>the root mean square error (RMS) of the event location</li> <li>the event identifier (evid) - begins with a “6” for events in the UUSS catalog and a “9” for new events</li> </ul> <p> </p>
NGC - Catalog
<p>NGC - Catalog from Carlos Toro and Sergio Barreiro Pérez for PRA1, matter Tipología y ciclo de vida de los datos</p>
The NGC Catalog, an exercise of data collection for UOC
<p>This project aims to download the NGC object catalog using webcrapping techniques and save it as a CSV file. The accessed URL is</p> <p><a href="https://in-the-sky.org/data/catalogue.php?cat=NGC&const=1&obj1Type=0&sort=0&view=1&page=x">https://in-the-sky.org/data/catalogue.php?cat=NGC&const=1&obj1Type=0&sort=0&view=1&page=x</a></p> <p>with 1 to 79 pages. The downloaded contents for each item compund the catalog are: Name = Name in the catalog type = Object kind mag = Visual magnitude of the object dist = Distance from the earth (when available) constellation = Constellation where to find the object RA = RA coordinates of the object. DEC = DEC Coordinates of the object str_name = Other name from the object. object_image = The image of the object.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.