Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,307

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,307 results for “libraries”

Learn how ShareScore rates datasets ↗
zenodo48/100

Science 2015 Farley et al Otx-a Library

<p>This dataset represents SEL-seq data generated by first constructing a dictionary of unique barcode tag-enhancer pairs by allowing 2 bp mismatches in the ~69 bp enhancers to buffer the sequencing mistakes. If more than one barcode tag was associated with a single enhancer the maximum reads between these tags were used. Barcoded tags that were attached to multiple enhancers were removed. The resulting dictionary contains 2,534,802 enhancers that are uniquely mapped to one or more barcode tags. 2 biological replicates were used in this experiment and the reads per million total reads (RPM) was used for each tag. In total 163,708 enhancers were detected by RNA-seq and 21,799 of them were defined as active enhancers by RPM &ge; 4 in either of the 2 replicates.</p> <p>Citation:&nbsp;Farley EK, Olson KM, Zhang W, Brandt AJ, Rokhsar DS, Levine MS. 2015. Suboptimization of developmental enhancers. <em>Science</em> <strong>350</strong>:325&ndash;328. doi:10.1126/science.aac6948</p> <p>Paper download link: https://www.science.org/doi/suppl/10.1126/science.aac6948/suppl_file/aac6948_tables_s1_and_s2.xlsx</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

S79 | UACCSCEC | Collision Cross Section (CCS) Library from UAntwerp

<p>This is the collection associated with list S79 UACCSCEC Collision Cross Section (CCS) Library from UAntwerp on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>A library containing the&nbsp;collision cross section (CCS) values&nbsp;of 311 adducts of 148 contaminants of emerging concern (CECs) and their metabolites&nbsp;measured with drift tube ion mobility high resolution mass spectrometry (in positive and negative ionization modes with&nbsp;N2 as drift gas) as described in Belova&nbsp;<em>et al.</em> (2021) DOI: <a href="https://pubs.acs.org/doi/10.1021/acs.analchem.1c00142">10.1021/acs.analchem.1c00142</a>.</p> <p>Changes: 27/04/2021 added new CIDs from deposition. 10/5/2021: added transformations table. 30/8/2022: corrected [M+H]+ for BDCIPP (CID <a href="https://pubchem.ncbi.nlm.nih.gov/compound/188119#section=Collision-Cross-Section">188119</a>) to 157.35 A^2 (from 178.72) upon request of the authors (see Belova <em>et al</em>. (2022) DOI: <a href="https://doi.org/10.1016/j.aca.2022.340361">10.1016/j.aca.2022.340361</a>).</p>

opencc-by-4.0Apr 2021View details →
zenodo48/100

Xylocopa sonorina - UCSB-IZC00012194 - Bee Library - 73e389aa-5886-4c48-8778-ba8932d1bd7e hash://sha256/96bfde1efa599e0e8e61de18b14d61dd308737f684950e4079c04e9bc0f33958 hash://md5/4940f68c84cffa4412f7ffb98bb255bd

<p>A biodiversity dataset graph: UCSB-IZC00012194</p> <p>The intended use of this archive is to facilitate (meta-)analysis of the Xylocopa sonorina - UCSB-IZC00012194 - Bee Library - 73e389aa-5886-4c48-8778-ba8932d1bd7e (UCSB-IZC00012194). UCSB-IZC00012194 provides an animated GIF, and wavefront 3D object model, of bee specimen Xylocopa sonorina UCSB-IZC00012194 University of Santa Barbara Invertebrate Zoology Collection as well as the original digital data/image files that were used to find and build this animated GIF and associated 3D model.&nbsp;</p> <p>This dataset provides versioned snapshots of the UCSB-IZC00012194 network as tracked by Preston [2,3] between 2022-09-26 and 2022-09-26 using &quot;preston update -u https://library.big-bee.net/portal/content/dwca/UCSB-IZC_DwC-A.zip&quot;.&nbsp;</p> <p>The archive consists of individual files with hexadecimal filenames (e.g., 03d2f9c6912935f54326d3e8c418cab6eddca5f69fb4f299e322cf2d114d0d03) to allow for parallel file downloads. The archive contains three types of files: index files, provenance logs and data files. Index files provide a way to links provenance files in time to establish a versioning mechanism. Provenance files describe how, when, what and where the UCSB-IZC00012194 content was retrieved. For more information, please visit https://preston.guoda.bio or https://doi.org/10.5281/zenodo.1410543 . &nbsp;</p> <p>To retrieve and verify the downloaded UCSB-IZC00012194 biodiversity dataset graph, download all files. Then, extract the archives into a &quot;data&quot; folder. Alternatively, you can use the preston[2] command-line tool to &quot;clone&quot; this dataset using:</p> <p>$ java -jar preston.jar clone --remote https://zenodo.org/record/7114321/files</p> <p>After that, verify the index of the archive by reproducing the following provenance log history:</p> <p>$ java -jar preston.jar history --log tsv<br> urn:uuid:0659a54f-b713-4f86-a917-5be166a14110&nbsp;&nbsp; &nbsp;http://purl.org/pav/hasVersion&nbsp;&nbsp; &nbsp;hash://sha256/9a5ab7b2278f2dea3fa329e9426dd4712e288b2586616e64986c8f09e76658c6&nbsp;&nbsp; &nbsp;<br> hash://sha256/af2bc3d2ac9ef865bedc33114e3c12232be58b46ac7033686d1eaed900c34a8d&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/9a5ab7b2278f2dea3fa329e9426dd4712e288b2586616e64986c8f09e76658c6&nbsp;&nbsp; &nbsp;<br> hash://sha256/781c17950a96d772c161552a8dff187ab427bcaa830f819758d0fdb8c60cf80e&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/af2bc3d2ac9ef865bedc33114e3c12232be58b46ac7033686d1eaed900c34a8d&nbsp;&nbsp; &nbsp;<br> hash://sha256/3239876613860452a47603946d3b961580447825b92cc7e566c3c9f4e8bb2b84&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/781c17950a96d772c161552a8dff187ab427bcaa830f819758d0fdb8c60cf80e&nbsp;&nbsp; &nbsp;<br> hash://sha256/419548ae006070af3ac9b1bbc90e9e9d51bf36131b23e8e057803cf7086a6842&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/3239876613860452a47603946d3b961580447825b92cc7e566c3c9f4e8bb2b84&nbsp;&nbsp; &nbsp;<br> hash://sha256/c9696c514e404b240d2362f0db545917e357d62c9f6c26b0b7a0b7df6444285e&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/419548ae006070af3ac9b1bbc90e9e9d51bf36131b23e8e057803cf7086a6842&nbsp;&nbsp; &nbsp;<br> hash://sha256/7d5a7fa413535390375687ff4cad53568b04761e5108da24400446bc7e8d57bb&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/c9696c514e404b240d2362f0db545917e357d62c9f6c26b0b7a0b7df6444285e&nbsp;&nbsp; &nbsp;<br> hash://sha256/86aec74994e16ea4bf509141b406546cdc491522705d948eee1f2b4ccefbd4b1&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/7d5a7fa413535390375687ff4cad53568b04761e5108da24400446bc7e8d57bb&nbsp;&nbsp; &nbsp;<br> hash://sha256/dcd61980ca9d78669e523fa643c9fa47e255481465384f98253f1b1ac7a5a8d0&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/86aec74994e16ea4bf509141b406546cdc491522705d948eee1f2b4ccefbd4b1&nbsp;&nbsp; &nbsp;<br> hash://sha256/8ecc7754cbab0c1ae169ed868bd9e4e68f7c592bb3609f27b32334c0a6c5e89f&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/dcd61980ca9d78669e523fa643c9fa47e255481465384f98253f1b1ac7a5a8d0&nbsp;&nbsp; &nbsp;<br> hash://sha256/d3c2c1ec6697a627607caab51135afa4b8d35c4795c9267f5c24ed3009b77fbe&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/8ecc7754cbab0c1ae169ed868bd9e4e68f7c592bb3609f27b32334c0a6c5e89f&nbsp;&nbsp; &nbsp;<br> hash://sha256/96bfde1efa599e0e8e61de18b14d61dd308737f684950e4079c04e9bc0f33958&nbsp;&nbsp; &nbsp;http://purl.org/pav/previousVersion&nbsp;&nbsp; &nbsp;hash://sha256/d3c2c1ec6697a627607caab51135afa4b8d35c4795c9267f5c24ed3009b77fbe&nbsp;&nbsp; &nbsp;</p> <p>To check the integrity of the extracted archive, confirm that each line produce by the command &quot;preston verify&quot; produces lines as shown below, with each line including &quot;CONTENT_PRESENT_VALID_HASH&quot;. Depending on hardware capacity, this may take a while.</p> <p>$ java -jar preston.jar verify<br> replace wwith preston verify | head -n4</p> <p>Note that a copy of the java program &quot;preston&quot;, preston.jar, is included in this publication. The program runs on java 8+ virtual machine using &quot;java -jar preston.jar&quot;, or in short &quot;preston&quot;.&nbsp;</p> <p>Files in this data publication:</p> <p>--- start of file descriptions ---</p> <p>-- description of archive and its contents (this file) --<br> README&nbsp;</p> <p>-- executable java jar containing preston [2,3] v0.4.5. --<br> preston.jar</p> <p>-- wavefront 3D object files<br> UCSB-IZC00012194.jpg<br> UCSB-IZC00012194.mtl<br> UCSB-IZC00012194.obj</p> <p>-- animated gifs<br> bee.gif<br> UCSB-IZC00012194.gif</p> <p>-- QR code<br> label.png</p> <p>-- preston archives containing UCSB-IZC00012194 data files, associated provenance logs and a provenance index --<br> 03d2f9c6912935f54326d3e8c418cab6eddca5f69fb4f299e322cf2d114d0d03<br> 064bc7772b2284c42917b785706ae72c196e9443dd394069ceee0f9cd8237e93<br> 093cbfe0e642bcea957785c6593db364ffe5254433bec3b3d0bb942764c47033<br> 0a921d571873916c6c806682e2ab8ead98213deb849f8fe72b34f7902174e655<br> 0d05a474f908c35350640aa006e11d27107b01cd8c6bcad020c474cb39f41dd2<br> 0d31c48fc6ac0556600d066f55c9eeac200ec0587f0af5a7a87de0768971b12c<br> 122f5df07bf185de58966e77522c3e8a1f7165e19c2a18cfa9e42952df1a4144<br> 13670f9ff787855776ab43fff4177871f922a86d96b579389b755e856d54b4b4<br> 13b28f9e8cd5b11ac419cbe8fef77ae35fee1ac83a3329ab9db0402645724564<br> 1710172af2410691c18d35e47550035c51ca8b0f3ce818bd66f48e61cdba1a7f<br> 181844ef7a5ff34460a19a5b82fa13abec38d531ba4420e57a5abe4ca3f5a69c<br> 19b4293d17d0e8205ef5e85e67c74ff4523665de333df2507a8a440e11b3ef88<br> 21d686d8d023c1d0d89447584106af0c920ca9d90f5f9a7116c2c584b8fa4416<br> 2281a46c916d9c559dfe56dd3e87ba23b196f8d7a9e42ccc00939015337a8c08<br> 290184df10c8a6390bc1debde079418e06dcef559d595f5d73fe0e9fff6e5879<br> 29285f2f9a7a1fe84c9ff6ee0117b102ad0dcf1530ee3bd42b5c70f9fd18a9d4<br> 2a5de79372318317a382ea9a2cef069780b852b01210ef59e06b640a3539cb5a<br> 2a9bed66ee211953ea7cba4c0b2e16022466cbc5696cc73e4d3db19f1a031541<br> 2bcd5566a650ca8b9156a4db514ed778770b02afa22a76b84f056ac042d857d7<br> 2cffff26667b33bfa79a8ae41177bd5c0ef25d9f60c15dddc32d613107bfc271<br> 2d90a1ba0fb369b680a3f20f128a457ce2e36d074b2aa8eb312a21c3d9b72b1a<br> 2ef8b9cced62d95a641c427300d063d0df51cc76be1c0538d1477694cab9f849<br> 2f125d3cbda2f34ea348965aa90ded47ec6868addd39c737c09856f8a768bdf3<br> 320d97685a0a0a8a4e3f6a610f732a65107d117c8a0870e1bb4982ccf6e59bad<br> 3239876613860452a47603946d3b961580447825b92cc7e566c3c9f4e8bb2b84<br> 32f069ef9c2de2a595aeed1ec69f3332df79e47e5f9c666ea0203d1948885a57<br> 33bd8fd545c1fd0b9708e0e51dd82cb33c59f74f655e66807d7f2da0761ef445<br> 35183865952ba1ce3788bec448f99ea718c758039580ca6d1966957cb7582514<br> 3c4e631a502de2a3f465043d4c9e0cce12d8b707da90b39a3768aa98651561df<br> 419548ae006070af3ac9b1bbc90e9e9d51bf36131b23e8e057803cf7086a6842<br> 453c012f50d765236d490c03f6b78ee740d16a0f6d183abc5746676cc5c78f3a<br> 46386e67d88207f7b98a103092cb5234c3aeafa9853244c69198f0aa78c18fc7<br> 47909127d1ac7821a3086c6c090d685bed8c1b07a7b9e7731df1f2caaf70cb35<br> 4c98d9ee8938c37c943245a8971b65c6cf8281e7526cf5d3205a6d5b19f537e9<br> 4d970ed5ece4889401091b3daaa74697a2b5123538bbc03159b5bebb8616e2da<br> 507eeebdafcf5211b249df63fcd14849cc9149c0621e8c4485f5c957f3e91445<br> 52ed56264e844af0c3e9850e09666034e596d1e2680815b818a4e474ccdb6a2a<br> 5447ddf194f37c246c89aefa699331e2072fc2b770789004e6a66a7d026de298<br> 5798de3b3434a6b1f8af9e0c57554e214ad1187c6c2479daefb3da1e53d8255f<br> 59179795306bce1502fbbe7e4d37bffc1f77f7f99b6b54685e3cc2c1e88c7b37<br> 5a6194035d3a780c882664b13bddbb280b08741400b506f617e8c95829efc129<br> 5b87eb636f6050d7fa6c07da1660729f708a48f4f6fd1f88e77d5c14c917342f<br> 5e4a90f2a7d51f22ab2ea9b669102a4265c33788f67aba7e37f0f17cef51ceae<br> 5eb0fe8feb8158002a62ad134d7ae25dd101d0389887172ac3351fdbc186e41d<br> 5f0c331eeedecaf4ad38b5fc8d833f92da1e24ab9ff45173c8b2cb05f4ed6f56<br> 600221bc9ce85dc3ba965883aeef24b5715637988f87b793681010912d1bdd21<br> 688dffb3a50c88200e791a83fa1fca215d26d8936fea003ba2be675c358702bd<br> 6e1f8a4f52112d3aac473ef6815c54b2dd11a1e097b8d68ea8f76950b7c78d24<br> 7703e881331f1402a72c8c2cafd5b206ea1548576353c5a3ea086d67b64f2a69<br> 773c7acb1b1f483cccacef2005bbb632f4ab06bc98fd771e9b859d84317c1770<br> 781c17950a96d772c161552a8dff187ab427bcaa830f819758d0fdb8c60cf80e<br> 786c8f13b48bf10228c3c404686eb3f262c422835efbade0b2956f5e5066e15d<br> 78919887ffc8206b3a9847ba6f7efd273c1f785145077277b3d8555eb6249475<br> 7a6e0e13e47e8cd92754a00ae4276d8718a18ef1dfb8965cec7790f4f9eb51dd<br> 7b9b77ce47ba5567b22eb6c93760d862b2227f51cc2d393406802a55d3670163<br> 7d5a7fa413535390375687ff4cad53568b04761e5108da24400446bc7e8d57bb<br> 81ceef6862891c899666acd4d75240555727a215dd4a57e1900ca94342383613<br> 849661f5418509327a6d3b180d267e49f152bd4d6553e055c83e911308f3d142<br> 8596fa69779a58fdc5d5079c427e0ae4f3f05fd59c097dbb3f4144bd2d042682<br> 8676b32841bc1aabf71c5e5830231864c0af29d0f59035e551f6900011b00c2e<br> 86aec74994e16ea4bf509141b406546cdc491522705d948eee1f2b4ccefbd4b1<br> 8e0ffe67171f86a2ca77f5a6d07709c7e39436fb317fbc059db1b854062ac441<br> 8ecc7754cbab0c1ae169ed868bd9e4e68f7c592bb3609f27b32334c0a6c5e89f<br> 96bfde1efa599e0e8e61de18b14d61dd308737f684950e4079c04e9bc0f33958<br> 983e3700682c637bb68ab8cbbe42dbfbe5268451881a0adc71092f926f783a2e<br> 988cf9b49358b074dbd712fc3a3af6eb0b043edc63446d66b625a419552e1d01<br> 9a5ab7b2278f2dea3fa329e9426dd4712e288b2586616e64986c8f09e76658c6<br> 9ba3244b9602020589ddd1bb96cb87ee7821908c4103c2ed8d3214803194f209<br> 9de90057f40f777d1ea071a3a7e98b55bfbfeb25d3a1e26c8beacfb32b9f09fa<br> 9e73e73c95c49c01d4f216e319c24e44ff34ae194b3c33edcd0654434bde7d30<br> a344330133331c5426dd588559cc1f491fac8b9a302ceaac659cddbd27c438cd<br> a41737eaea1ac2948da4f7baca15c70628dd7930613b962c4b2054f20ba6442f<br> af2bc3d2ac9ef865bedc33114e3c12232be58b46ac7033686d1eaed900c34a8d<br> b0e3553c65bbdfc3e0f46addd71eb482fe2100e875bc162b6f4729af69e6e2be<br> b391b0d46818ae623c7d44430db18e1b4b61f206f6b63ec36c6712b5be9a38e8<br> b60bc829327a5d0e4202abde829023529d8a4c75116f2bed7b775f26e92fbc61<br> c3bbf1f8c4ef1f564ddb59c10c9d597e7a49e8399a023f39f20e73bdf0fb87b8<br> c58be26ee6726a477936ec1df0795700abfb9a6a9c6365169145247a67a0b787<br> c92fb4886172ef60f95d77c4bdc3e762ce36d501f5fb78a23402b0134f0bcce9<br> c9696c514e404b240d2362f0db545917e357d62c9f6c26b0b7a0b7df6444285e<br> d3c2c1ec6697a627607caab51135afa4b8d35c4795c9267f5c24ed3009b77fbe<br> d60647ca1ffd0aaeb60b0467c6676a48d1d254ee47af63349d9e7994572afe67<br> dcd61980ca9d78669e523fa643c9fa47e255481465384f98253f1b1ac7a5a8d0<br> e0ec9708f3a418423166c71cabb0520c34293db024d1ea5e100924eb6d6337be<br> e1ce89711faa2a81e40f38560166da4ac36c51d17a467705e559d6b3f4c0bf31<br> e29883205643e1eff6efd121ae912cf4a5b96c5301585c118acfc0f7571ad980<br> e5ffada14e04ec2b521b2239e0d183ddcf6feb76e88810ff700a6c9cac309c9e<br> ec564b934f9524c3c92e625eac3709afa5f010d62fa4b879c5a9704b49f5fcbd<br> ec6bf13afb42287178591a04f7009f5d822da9c62e34e9bdf9b4d6b7bbb8b129<br> ed6aff81c806670d4650476a59dc22fb88034d610cfc6769c5deb9d26a810de9<br> f446aef17a59caac1f0571eba1632e1e17747ecf6d87c3dc99d8f54c4c549932<br> f50e03ae29d11e4f895d8df7fbe90f6c7c221f9e932048c51adf9e0229c97884<br> f8d2bc8175771d210e911268feb7746a18e9b0d3d303f04ae469b4d5059bc90f<br> fe14ffe132b80a4ea5724ae376fe1a4aebcf54259487c736928063fd7d0c3e6b&nbsp;</p> <p>--- end of file descriptions ---</p> <p><br> References&nbsp;</p> <p>[1] Xylocopa sonorina - UCSB-IZC00012194 - Bee Library - 73e389aa-5886-4c48-8778-ba8932d1bd7e (UCSB-IZC00012194, https://library.big-bee.net/portal/content/dwca/UCSB-IZC_DwC-A.zip) accessed from 2022-09-26 to 2022-09-26 with provenance hash://sha256/96bfde1efa599e0e8e61de18b14d61dd308737f684950e4079c04e9bc0f33958.<br> [2] https://preston.guoda.bio, https://doi.org/10.5281/zenodo.1410543 .&nbsp;<br> [3] MJ Elliott, JH Poelen, JAB Fortes (2020). Toward Reliable Biodiversity Dataset References. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2020.101132</p> <p>This project made possible by National Science Foundation Awards: 1839201, 2102006, 2101929, 2101908, 2101876, 2101875, 2101851, 2101345, 2101913, 2101891 and 2101850.</p>

opencc-zeroSep 2022View details →
zenodo48/100

Diachronic word embeddings from 19th-century newspapers digitised by the British Library (1800-1919)

<p>Word vectors related to the paper&nbsp;<em>Machines in the media: semantic change in the lexicon&nbsp;of mechanization in 19th-century British newspapers&nbsp;</em>by Nilo Pedrazzini and Barbara McGillivray (2022).</p> <p>The embeddings were trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 1 window = 3 vector_size = 200 epochs = 5</code></pre> <p>The embeddings&nbsp;are divided into periods of ten years each, with the vectors from each decade aligned to the ones from the most recent decade (1910s) using Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project webpage (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Near-infrared (NIR) soil spectral library using the NeoSpectra Handheld NIR Analyzer by Si-Ware

<p>Up-to-date information on soil properties and the ability to track changes in soil properties over time are critical for improving multiple decisions on soil security at various scales, ranging from global climate change modeling and policy to national level environmental and development planning, to farm and field level resource management. Diffuse reflectance infrared spectroscopy has become an indispensable laboratory tool for the rapid estimation of numerous soil properties to support various soil mapping, soil monitoring, and soil testing applications. Recent advances in hardware technology have enabled the development of handheld sensors with similar performance specifications as laboratory-grade near-infrared (NIR) spectrometers.</p> <p>Here, we've compiled a hand-held NIR spectral library (1350-2550 nm) using the NeoSpectra Handheld NIR Analyzer developed by <a href="https://www.si-ware.com/">Si-Ware</a>. Each scanner is fitted with Fourier-Transform technology based on the semiconductor Micro Electromechanical Systems (MEMS) manufacturing technique, promising accuracy, and consistency between devices.</p> <p>This library includes 2,106 distinct mineral soil samples scanned across 9 of these portable low-cost NIR spectrometers (indicated by serial no). 2,016 of these soil samples were selected to represent the diversity of mineral soils found in the United States, and 90 samples were selected across Ghana, Kenya, and Nigeria. 519 of the US samples were selected and scanned by <a href="https://www.woodwellclimate.org/">Woodwell Climate Research Center</a>. These samples were queried from the <a href="https://ncsslabdatamart.sc.egov.usda.gov/">USDA NRCS NSSC-KSSL Soil Archives</a> as having a complete set of eight measured properties (TC, OC, TN, CEC, pH, clay, sand, and silt). They were stratified based on the major horizon and taxonomic order, omitting the categories with less than 500 samples. Three percent of each stratum (i.e., a combination of major horizon and taxonomic order) was then randomly selected as the final subset retrieved from KSSL's physical soil archive as 2-mm sieved samples. The remaining 1,604 US samples were queried from the USDA NRCS NSSC-KSSL Soil Archives by the <a href="https://www.unl.edu/">University of Nebraska - Lincoln</a> to meet the following criteria: Lower depth &lt;= 30 cm, pH range 4.0 to 9.5, Organic carbon &lt;10%, Greater than lower detection limits, Actual physical samples available in the archive, Samples collected and analyzed from 2001 onwards, Samples having complete analyses for high-priority properties (Sand, Silt, Clay, CEC, Exchangeable Ca, Exchangeable Mg, Exchangeable K, Exchangeable Na, CaCO3, OC, TN), &amp; MIR scanned.</p> <p>All samples were scanned dry 2mm sieved. ~20g of sample was added to a plastic weighing boat where the NeoSpectra scanner would be placed down to make direct contact with the soil surface. The scanner was gently moved across the surface of the sample as 6 replicate scans were taken. These replicates were then averaged so that there is one spectra per sample per scanner in the resulting database.</p> <p>A subset of 1,976 US topsoil samples was used to create Cubist models for 8 soil properties including bulk density (BD, &lt;2mm fraction, 1/3 Bar, units in grams per cubic centimeter), calcium carbonate (CaCO3, &lt;2mm fraction, units in weight percent), clay content (percent), buffered ammonium-acetate exchangeable potassium (Ex. K, units in centimoles of charge per kilogram of soil), pH, sand content (percent), silt content (percent), and estimated organic carbon (SOC, estimated after inorganic carbon removal, units in weight percent). Two strategies were evaluated for handling scanner-to-scanner variability: averaging scans per sample (avg) versus retaining replicate scans across all scanners (reps) during model building. Cubist avg models and cubist reps models are provided here for the 8 soil properties outlined in &ldquo;.qs&rdquo; file format and can be opened and worked with in the R programming language. The subset of 1,976 samples has also been provided here for reproducibility (1976_NSlibrary_withmetadata.csv).</p> <p>The&nbsp;repository contains:</p> <ul> <li><em>Neospectra_database_column_names.csv</em>: describes the variables (columns) of site and soil data, and the range of near-infrared (NIR, 1350-2550 nm) and mid-infrared (MIR, 600-4000 cm-1) spectra. The CSV is composed of the file name, column name, type, example, and description with measurement unit.</li> <li><em>Neospectra_WoodwellKSSL_MIR.csv</em>: the equivalent MIR spectra of neospectra samples fetched from the KSSL database and formatted to the OSSL specifications.</li> <li><em>Neospectra_WoodwellKSSL_soil+site+NIR.csv</em>:&nbsp;soil, site, and Neospectra's NIR. Each row&nbsp;contains one&nbsp;replicated spectra of a given scanner (6 repeats per scanner per soil sample). Soil and site info is filled within the same soil sample.</li> <li>1976_NSlibrary_withmetadata.csv: data matrix for reproducible model calibration.</li> <li>Models: <ul> <li>log..bd_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for log(1+BD).</li> <li> <p>log..caco3_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for log(1+CaCO3).</p> </li> <li> <p>clay_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for clay.</p> </li> <li> <p>log..k.ex_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for log(1+Ex. K).</p> </li> <li> <p>ph.h2o_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for pH.</p> </li> <li> <p>sand_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for sand.</p> </li> <li> <p>silt_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for silt.</p> </li> <li> <p>log..soc_model_nir.neospectra_cubist_AVG_ossl_na_v1.2.qs: Cubist average NIR model for log(1+SOC).</p> </li> <li> <p>log..bd_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: &nbsp;Cubist replicates NIR model for log(1+BD).</p> </li> <li> <p>log..caco3_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for log(1+CaCO3).</p> </li> <li> <p>clay_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for clay.</p> </li> <li> <p>log..k.ex_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for log(1+Ex. K).</p> </li> <li> <p>ph.h2o_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for pH.</p> </li> <li> <p>sand_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for sand.</p> </li> <li> <p>silt_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for silt.</p> </li> <li> <p>log..soc_model_nir.neospectra_cubist_REPS_ossl_na_v1.2.qs: Cubist replicates NIR model for log(1+SOC).</p> </li> </ul> </li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Materials for 2d representation of the HathiTrust Library

<p>Materials to create the LargeVis visualization online at&nbsp;http://creatingdata.us/datasets/hathi-features/, and described in&nbsp;<em>Benjamin Schmidt, &quot;Stable random projection: lightweight, general-purpose dimensionality reduction for digitized libraries,&quot; Journal of Cultural Analytics. October 3, 2018.</em></p> <p>Two items. First, `hathi_pca.bin`: a binary file with 100-dimensional representations of the complete Hathi Trust Extended Features set. These began as 1280-dimensional SRP features, and were reduced to 100 dimensions using a PCA transformation matrix derived using a random sample of the full 13 million book set. Vectors were reduced to unit length before PCA, but not afterwords; this means that in general, their length gives some sense of much information was lost in the PCA representation. This can be read using the code at https://github.com/bmschmidt/pySRP, or anything that reads word2vec formatted vectors. Includes HathiTrust identifiers.</p> <p>Second, `hathi.tsv.gz`: a row oriented set containing a variety of metadata fields for each set, including (as &#39;x&#39; and &#39;y&#39;) the coordinates of a 2-d LargeVis visualization. This is the immediate input to the visualization at&nbsp;ttp://creatingdata.us/datasets/hathi-features/. Columns should be relatively straightforward; they are derived from the HathiTrust MARC records, which can be accessed through Hathi&#39;s public API. Classification codes (&#39;lc1&#39;) are using the Library of Congress classification; they represent the subclass (generally two characters, though it can be one or three). The first character alone represents the LC class and can be useful for coloring high-level overviews.</p> <p>These two files can be merged through the Hathi Trust identifier present in both.</p> <p>&nbsp;</p>

opencc-by-sa-4.0Nov 2018View details →
zenodo48/100

An acoustically isolated European starling song library

<p>A dataset of song collected from 14 European starlings individually recorded in acoustically isolated chambers.&nbsp;Each folder contains&nbsp;vocalizations for one bird.</p> <p>These&nbsp;data were&nbsp;used for the publication, &quot;<em>Parallels in the sequential organization of birdsong and human speech</em>&quot;. Nature Communications (2019). If you use this dataset, please cite this publication and this repository.</p> <p>Work supported by NSF Graduate Research Fellowship 2017216247 to TS and an&nbsp;NIH R56DC016408&nbsp;to TQG.</p>

opencc-by-4.0Jun 2019View details →
zenodo48/100

Berlin State Library (2024). Metadata of the Digitized Collections of the Berlin State Library (SBB)

<p>The motivation for creating this dataset was to enable research on the basis of metadata which are available in a cultural heritage institution on a large scale. Libraries such as the Staatsbibliothek zu Berlin &ndash; Berlin State Library (SBB) typically provide three kinds of data: Images (scans of books, illustrations contained in the scanned material, or else), texts (OCR'd from digitized books or manuscripts), and metadata. However, metadata form an underresearched resource, which is lamentable: These metadata are of a high quality since they have been established by trained librarians, archivists, or other cultural heritage practitioners. The publication of a set of metadata of more than 200.000 works aims therefore at providing an underresearched high-quality type of data. The basic interest of the funder in this data publication is the stimulation of innovation.</p> <p>The dataset consists of a single table containing the metadata of all 219.419 works which were available in the Digitized Collections of the Berlin State Library (SBB) on July 29th, 2024. The size of the .parquet file is about 46 MB.</p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

LEUKOS' dataset: A selection of sixteenth and seventeenth century Stambøger from the Royal Danish Library (Copenhagen)

<p>The data were generated to document LEUKOS&rsquo; research for objective 1, point 3 (O1.3) of the research project. LEUKOS&rsquo; research question (RQ) and brief description of O1: LEUKOS investigates whether and how, during his stay in Wittenberg (1586-1588), Bruno&rsquo;s notions of the soul and of language were influenced by the new Protestant understandings of philosophy as evidenced by their discussions and suggested reforms of the liberal arts at the philosophical faculty and in Melanchthon&rsquo;s works, and drew conclusions regarding the relation between Giordano Bruno and the Reformation. To reconstruct the social, intellectual, and political context in Wittenberg during Bruno&rsquo;s stay (1586-1588) and pave the way for the interpretation of Bruno&rsquo;s praise of the Reformers as tolerant (O1), LEUKOS reconstructs the environment that Bruno joined in Wittenberg at the time when Aristotle was reintroduced in the curricula of Luther&rsquo;s university by focusing (O1.1) on the debate concerning the soul in Wittenberg (Gnesio-Lutherans and Philippists) in its philosophical and theological arguments, but also by referring to the political implications of this debate in the view of tolerance; (O2.2) on the role of female theologians and intellectuals in the reformed communities, and in the intellectual world, as the Reformation gave them access to direct participation in cultural and entrepreneurial professions (e.g. as printers); (O1.3) on Danish intellectuals in Bruno&rsquo;s network as Wittenberg was a centre of diffusion of the Reformation also for the Scandinavian countries. Data generated: pictures of a selection of sixteenth and seventeenth century Stamb&oslash;ger (Alba Amicorum) belonging to the Royal Danish Library collections documenting the transnational intellectual network between Germany and Copenhagen in the second half of the sixteenth century and in the first half of the seventeenth century. The data might be useful to researchers working (1) on the history of the book, (2) on the Royal Danish Library sixteenth and seventeenth century book collections, (3) on Stamb&oslash;ger (Alba Amicorum), (4) on intellectual networks between Germany and Denmark in the sixteenth and seventeenth century.</p> <p>Instrument- or software-specific information needed to interpret the data: It is possible to open a HEIC file on Windows 10 or 11 by: 1. Using the Microsoft Photos App: If your Windows 10 or 11 is up to date, the Microsoft Photos app should already support the HEIC file format. Simply right-click on the HEIC file, select "Open With," and choose "Photos" from the list of apps. The Microsoft Photos app will open the HEIC file and display its contents; 2. Converting HEIC files to JPEG or PNG: If the Microsoft Photos app doesn't support HEIC files on your system, you can convert HEIC files to JPEG or PNG format using an online converter or dedicated software. Search for "HEIC to JPEG converter" or "HEIC to PNG converter" in your preferred search engine to find various online converter tools or downloadable software. Upload your HEIC file to the converter tool or software, select the desired output format (JPEG or PNG), and convert the file. Once converted, you can easily view the resulting JPEG or PNG file using any image viewer or photo app on your Windows 10 or 11 PC; 3. Installing a Third-party HEIC Viewer: If you frequently work with HEIC files, you can also install a dedicated HEIC viewer app from the Microsoft Store or other reputable sources. Search for "HEIC viewer" in the Microsoft Store or other platforms, and look for apps that specifically support viewing HEIC files. Install the chosen HEIC viewer app, and use it to open and view HEIC files on your Windows 10 or 11 PC.</p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Spectral library of vegetation from Mediterranean woodlands

<p>Site description:</p> <p>All reflectance measurements have been collected in Mediterranean oak woodland at <em>Herdade </em>da <em>Machoqueira do Grou</em>, located<em>&nbsp;</em>in Central Portugal (39&deg; 08&prime; 18.9&Prime; N, 9&deg; 19&prime; 56.22&Prime; W, 165-m height). The site is characterized by a Mediterranean climate with mild winters and hot dry summers. The average annual precipitation recorded at the climate station of Santar&eacute;m (39&deg; 12&prime; N, 8&deg; 44&prime; W) for the period 1981&ndash;2010 was 652 mm, and mean daily temperature was&nbsp; 17&deg;C&nbsp; (<a href="http://www.ipma.pt/pt/oclima/normais.clima/">www.ipma.pt/pt/oclima/normais.clima/</a>). Detailed meteorological measurements of radiation, temperature, and air humidity are also publicly available (Cerasoli et al., 2020). The soil is a cambisol (FAO) with 81% sand, 5% clay, and 14% silt. The tree layer is represented exclusively by cork oak trees (<em>Quercus suber</em> L.) with a tree density of 177 tree ha<sup>-1</sup> and leaf area index (LAI) of 1.5. The mean total tree height and height below the canopy are 7.9 and 3.1m respectively (Cerasoli et al., 2015). Tree canopy represents 36% of the soil cover fraction. The understorey is composed of a mixture of shrubs and herbaceous species. The site was plowed in 2013 (Correia et al., 2016), hence the cover fraction of shrubs changed across years. A field survey in 2017 estimated an 18% coverage of shrubs and 41% of herbaceous species, while the remaining 41% was represented by litter and bare soil (Heuschmidt et al., 2020). The most represented shrub species are <em>Cistus salvifolius</em> (cistus) and the <em>Ulex airensis</em> (ulex). In spite of occupying the same habitat, the two species have different growth habits and stress strategies. While the cistus is a semi-deciduous species with shallow roots, decreasing its canopy area during the summer period, the ulex has a deep root system and spine shaped leaves and shoots conferring high drought resistance (Correia et al., 2014). The herbaceous layer is composed of C3 species mainly grasses (44.5%) and legumes (28.7%) (Cerasoli et al., 2015).</p> <p>&nbsp;</p> <p>Reflectance measurements:</p> <p>All spectral observations were acquired with an ASD FieldSpec3 spectroradiometer (Malvern Panalytical, Boulder, USA) in the range of 350-2300nm. The visible and near-infrared region (350-1000nm) has a spectral resolution (full-width half maximum) of 3nm and a sampling interval of 1.4nm, while the mid infrared region (1000-2500nm) has a spectral resolution of 10nm and a sampling interval of 2.0nm. Canopy spectral data were collected by a fiber optic cable inserted into a pistol grip. A white reference of known reflectance (Spectralon panel, Labsphere, Inc., North Sutton, USA) was used to normalize for variation in atmospheric conditions and to convert the measurements into absolute reflectance. All targets were fully exposed to solar radiation at the time of the measurements. Measurements were performed on cork oak, cistus, and ulex canopies. Herbaceous plots were delimited by a 50X50 cm quadrat. Oak trees canopy measurements were done using a scaffold on the south side of the canopy. All canopy measurements were performed with a nadir view, a field of view angle of 25&ordm;, and a distance of about 90cm from the target, which resulted in a field of view of about 1256 cm<sup>2</sup>. All spectra were collected for 2 hours around solar noon, to minimize the effects of shadowing and solar zenith changes, with five replicates for each target, representing each the average of 25 spectra. All reflectance values in the range 1350-1400nm and 1800-1950nm were excluded, corresponding to the atmospheric water vapor absorption regions. A leaf clip including a white and a black standard was used for the measurement of the reflectance of cork oak leaf blades avoiding main veins.</p> <p>&nbsp;</p> <p>File description:&nbsp;</p> <p>The file &quot;specveg_data_spectra&quot; concerns all spectral data, the &quot;specveg_metadata&quot; covers the additional data of every single measured vegetation including&nbsp;photos (URL),&nbsp;and&nbsp;the &quot;specveg_meta&quot; describes all the existing variables.</p>

opencc-by-4.0Aug 2021View details →
zenodo48/100

GNPS Drug Library spectral files and metadata

<p>Global Natural Product Social Molecular Networking (GNPS) Drug Library: a centralized collection of reference spectra for drugs and their metabolites/analogs along with structured pharmacologic metadata including exposure source, pharmacologic class, therapeutic indication, and mechanism of action.&nbsp;</p> <p>Two MS/MS reference libraries:&nbsp;</p> <ul> <li>GNPS_Drug_Library_Spectra_Drugs_and_Metabolites.mgf: Reference spectra for drugs and drug metabolites collected from the GNPS Spectral Library and MSnLib.</li> <li>GNPS_Drug_Library_Spectra_Drug_Analogs.mgf: MS/MS spectra analogs of drugs in publicly accessible untargeted metabolomics data derived from spectral alignment strategies.</li> </ul> <p>Two metadata files:</p> <ul> <li><span>GNPS_Drug_Library_Metadata_Drugs.csv: Controlled-vocabulary pharmacologic metadata on the drugs.&nbsp;</span></li> <li><span>GNPS_Drug_Library_Metadata_Drug_Analogs.csv: Metadata on the drug analogs, including connections to parent drugs, mass offsets, and pharmacologic metadata based on the parent drugs.</span></li> </ul> <p><span>Publication:</span></p> <p><span>https://www.biorxiv.org/content/10.1101/2024.10.07.617109v1</span></p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB) Version 2 - August 2025

<p>This dataset was created with the intent to provide a single larger set of metadata from Berlin State Library for research purposes and the development of AI applications.</p> <p>The dataset comprises descriptive metadata of 2.639.554 titles derived from the union catalogue K10plus, a database with about 200 million records from libraries across 11 German states. Selected are all records that include system entries ("Systemstellen") from the historical classification of the "Alter Realkatalog" (ARK), a subject catalogue of the Staatsbibliothek zu Berlin &ndash; Berlin State Library. They refer to publications from 1501 to 1955 and reproductions thereof. The title data contain subject headings and BK classmarks that have been transmitted onto them from the "<a href="https://ark.staatsbibliothek-berlin.de/">Historische Systematik</a>", the online representation of the ARK classification.</p>

opencc-zeroJul 2024View details →
zenodo48/100

Open Soil Spectral Library (training data and calibration models)

<p><strong>Open Soil Spectral Library</strong> contains training MIR (91,631) and VisNIR (65,063) spectral scans + soil calibration data (&gt;60,000 unique locations) and calibration models. Key data set:</p> <ul> <li>ossl_all_L1_v1.2.qs: soil laboratory, site and spectra information;</li> </ul> <p>Important note: The data set spatially over-represents USA and European Union, with little training data in Asia, South America and Australia, hence calibration models reflect primarily soils of USA and Europe.</p> <p>To use the models and data please install <a href="https://hub.docker.com/r/opengeohub/r-geo">R and required packages</a>. Read more about the <strong><a href="https://github.com/traversc/qs">QS data format</a></strong> and how to convert it to CSV or similar. Modeling steps are explained in detail in: <a href="https://github.com/soilspectroscopy/ossl-models">https://github.com/soilspectroscopy/ossl-models</a>. To visualize database please use: <a href="https://explorer.soilspectroscopy.org/">https://explorer.soilspectroscopy.org/</a></p> <p>Complete OSSL documentation can be found at: <a href="https://soilspectroscopy.github.io/ossl-manual/">https://soilspectroscopy.github.io/ossl-manual/</a></p> <p><a href="https://soilspectroscopy.org/"><strong>Soil Spectroscopy for the Global Good</strong></a> is a Coordinated Innovation Network funded by USDA NIFA Food and Agriculture Cyberinformatics Tools Program (<a href="https://nifa.usda.gov/press-release/nifa-invests-over-7-million-big-data-artificial-intelligence-and-other">Award #2020-67021-32467</a>).</p> <p>Input datasets are property of the <a href="https://www.nrcs.usda.gov/wps/portal/nrcs/main/soils/research">USDA NRCS National Soil Survey Center &ndash; Kellogg Soil Survey Laboratory</a>, <a href="https://www.worldagroforestry.org/">ICRAF-World Agroforestry</a>, <a href="https://www.isric.org/">ISRIC-World Soil Information</a>, the <a href="http://africasoils.net/services/data/soil-databases/">Africa Soil Information Service</a> funded by the Bill and Melinda Gates Foundation, the <a href="https://esdac.jrc.ec.europa.eu/">European Soil Data Centre</a>, the <a href="https://www.neonscience.org/">National Ecological Observatory Network</a>, and <a href="https://sae.ethz.ch/">ETH Zurich</a>.&nbsp;</p> <p>For more advanced uses of the soil spectral libraries <strong>we advise to contact the original data producers</strong> especially to get help with using, extending and improving the original SSL data.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

UniCat library catalogue dataset

<p>This dataset was created in the context of a project about libcitations or library catalogue analysis of book publications by Flemish Social Sciences and Humanities researchers.&nbsp;</p> <p>The dataset was constructed on the basis of&nbsp;the openly available API of the Belgian UniCat library catalogue (<a href="https://www.unicat.be/uniCat?func=search&amp;uiLanguage=en">UniCat-Search</a>). The library catalogue was searched in September 2021 by a matching of the catalogue against the ISBNs present in the VABB-SHW database (see&nbsp;Aspeslagh, Peter, Guns, Raf, &amp; Engels, Tim C.E. (2021). VABB-SHW: Dataset of Flemish Academic Bibliography for the Social Sciences and Humanities (edition 11) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5795899).</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Decade-level Word2Vec models from automatically transcribed 19th-century newspapers digitised by the British Library (1800-1919)

<p>Word embeddings trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 5 window = 5 vector_size = 100 epochs = 5</code></pre> <p>The embeddings&nbsp;are divided into periods of ten years each. Unlike those in <a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, these were not aligned and OCR errors skimmed from the vocabulary.&nbsp;</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project website (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>

opencc-by-4.0May 2023View details →
zenodo48/100

FAIR-CHO Citation Model Zotero Group Library Bibliography. Supplementary material

<p>This is a selected bibliography created during the project <em>A FAIR-enabling citation model for Cultural Heritage Objects</em>.</p> <p>This bibliography has been set up via a Zotero Library Group, by organizing it into subject folders representative of the project content. An initial list of descriptors was also defined to &#39;semantically&#39; label the bibliographic references as they were collected.</p> <p>The dataset represents the Zotero Library on 1st August 2023.</p> <p>The dataset is published in .csv and .ris formats.</p> <p>See also on Zotero Groups: <a href="https://www.zotero.org/groups/4883319/cho_citation_model/library">https://www.zotero.org/groups/4883319/cho_citation_model/library</a>.</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

The Brazilian Soil Spectral Library (VIS-NIR-SWIR-MIR) Database: Open Access

<p><strong>Abstract:</strong></p> <p>NEW VERSION V.002 (Some Lat Long Coordinates added).</p> <p>Soil spectroscopy has emerged as a solution to the limitations associated with traditional soil surveying and analysis methods, addressing the challenges of time and financial resources. Analyzing the soil&#39;s spectral reflectance enables to observe the soil composition and simultaneously evaluate several attributes because the matter, when exposed to electromagnetic energy, leaves a &quot;spectral signature&quot; that makes such evaluations possible. The Soil Spectral Library (SSL) consolidates soil spectral patterns from a specific location, facilitating accurate modeling and reducing time, cost, chemical products, and waste in surveying and mapping processes. Therefore, an open access SSL benefits society by providing a fine collection of free data for multiple applications for both research and commercial use.</p> <p><strong>BSSL Description and Usefulness</strong></p> <p>The Brazilian Soil Spectral Library (BSSL), available at&nbsp;<a href="https://bibliotecaespectral.wixsite.com/english">https://bibliotecaespectral.wixsite.com/english</a>, is a comprehensive repository of soil spectral data. Coordinated by JAM Dematt&ecirc; and managed by the GeoCiS research group, the BSSL was initiated in 1995 and published by Dematt&ecirc; and collaborators in 2019. This initiative stands out due to its coverage of diverse soil types, given Brazil&#39;s significance in the agricultural and environmental domains and its status as the fifth largest territory in the world (IBGE, 2023). In addition, a Middle Infrared (MIR) dataset has been published (Mendes et al., 2022), part of which is included in this repository. The database covers 16,084 sites and includes harmonized physicochemical and spectral (Vis-NIR-SWIR and MIR range) soil data from various sources at 0-20 cm depth.&nbsp;All soil samples have Vis-NIR-SWIR data, but not all have MIR data.</p> <p>The BSSL provides open and free access to curated data for the scientific community and interested individuals. Unrestricted access to the BSSL supports researchers in validating their results by comparing measured data with predicted values. This initiative also facilitates the development of new models and the improvement of existing ones. Moreover, users can employ the library to test new models and extract information about previously unknown soil properties. With its extensive coverage of tropical soil classes, the BSSL is considered one of the most significant soil spectral libraries worldwide, with 42 institutions and 61 researchers participating. However, 47 collaborators from 29 institutions have authorized the data opening. Other researchers can also provide their data upon request through the coordinator of this initiative.</p> <p>The data from the BSSL project can also help wet labs to improve their analytical capabilities, contributing to developing hybrid wet soil laboratory techniques and digital soil maps while informing decision-makers in formulating conservation and land use policies. The soil&#39;s capacity for different land uses promotes soil health and sustainability.</p> <p><strong>Coverage</strong></p> <p>The BSSL data covers all regions of Brazil, including 26 states and the Federal District. It is in a&nbsp;<em>.xlsx</em>&nbsp;format and has a total size of 305&nbsp;Mb. The table is structured in sheets with rows for observations,&nbsp;and columns,&nbsp;representing various soil attributes in the surface layer, from 0 to 20 cm depth. The database includes environmental and physicochemical properties (22 columns and 16,084 rows), Vis-NIR-SWIR spectral bands (2151 columns and 16,084 rows), and MIR channels (681 columns and 1783 rows). An ID unique column can merge the sheet for each attribute or spectral range.</p> <p><strong>Accessing original data source</strong></p> <p>Using these data requires their reference in any situation under copyright infringement penalty. Three mechanisms are available for users to reach the original and complete data contributors:</p> <p>a) Refer to sheet two for name and code-based searches;</p> <p>b) Visit the website&nbsp;<a href="https://bibliotecaespectral.wixsite.com/english/lista-de-cedentes">https://bibliotecaespectral.wixsite.com/english/lista-de-cedentes</a>&nbsp;or locate the contributors&#39; list by Brazilian state;</p> <p>c) Visit the website of the Brazilian Soil Spectral Service &ndash; Braspecs <a href="http://www.besbbr.com.br/">http://www.besbbr.com.br/</a>, an online platform for soil analysis that uses part of the current SSL (Dematt&ecirc; et al., 2022) - It was developed and managed by GeoCiS. There, owners from all over the country can be found.</p> <p><strong>Proceeding to data analysis</strong></p> <p>We registered and organized the samples at the ESALQ/USP Soil Laboratory. Some samples arrived without preliminary data analyses, so we analyzed them for soil organic matter (SOM), granulometry, cation exchange capacity (CEC), pH in water, and the presence of Ca, Mg, and Na, following the recommendations of Donagemma et al. (2011).</p> <p>The GeoCiS research group performed spectral analyses following the procedures described by Bellinaso et al. (2010). Dematt&ecirc; et al. (2019) provide detailed methods for sampling, preparation, and soil analyses, including reflectance spectroscopy.&nbsp;Latitude and longitude data can be requested directly from the data owner.&nbsp;In summary, the following steps are involved in data acquisition.</p> <p>a) We subjected the soil samples to a preliminary treatment, which involved drying them in an oven at 45&deg;C for 48 hours, grinding them, and sieving them through a 2mm mesh;</p> <p>b) We placed the samples in Petri dishes with a diameter of 9 cm and a height of 1.5 cm;</p> <p>c) We homogenized and flattened the surface of the samples to reduce the shading caused by larger particles or foreign bodies, making them ready for spectral readings;</p> <p>d) The spectral analyses took place in a darkened room to avoid interference from natural light. We used a computer to record the electromagnetic pulses through an optical fiber connected to the sensor, capturing the spectral response of the soil sample;</p> <p>e) We obtained reflectance data in the Visible-Near Infrared-Shortwave Infrared (Vis-NIR-SWIR) range using a FieldSpec 3 spectroradiometer (Analytical Spectral Devices, ASD, Boulder, CO), which operates in the spectral range from 350 to 2500 nm;</p> <p>f) The sensor had a spectral resolution of 3 nm from 350-700 nm and 10 nm from 700-2500 nm, automatically interpolated to 1 nm spectral resolution in the output data, resulting in 2151 channels (or bands); and</p> <p>g) We positioned the lamps at 90&deg; from each other and 35 cm away from the sample, with a zenith angle of 30&deg;.</p> <p>The sensor captured the light reflected through the fiber optic cable, which was positioned 8 cm from the sample&#39;s surface.</p> <p>We used two 50W halogen lamps as the power source for the artificial light. It&#39;s important to note that we took three readings for each sample at different positions by rotating the Petri dish by 90&deg;.</p> <p>Each reading represents the average of 100 scans taken by the sensor. From these three readings, we calculated the final spectrum of the samples. Notably, the laboratory&#39;s equipment and procedures for soil sample spectral analyses followed the ASD&#39;s recommendations, particularly about sensor calibration using a white spectralon plate as a 100% reflectance standard.</p> <p>For the analysis in the Middle Infrared (MIR) spectral region, we followed the procedures outlined by&nbsp;Mendes et al. (2022). We milled the soil fraction smaller than 2 mm, sieved it to 0.149 mm, and scanned it using a Fourier Transform Infrared (FT-IR) alpha spectroradiometer (Bruker Optics Corporation, Billerica, MA 01821, USA) equipped with a DRIFT accessory.</p> <p>The spectroradiometer measured the diffuse reflectance using Fourier transformation in the spectral range from 4000 cm<sup>-1</sup>&nbsp;to 600 cm<sup>-1</sup>, with a resolution of 2 cm<sup>-1</sup>. We conducted these measurements in the Geotechnology Laboratory of the Department of Soil Science at Esalq-USP. We took the average of 32 successive readings to obtain a soil spectrum. Sensor calibration took place before each spectral acquisition of the sample set by standardizing it against the maximum reflectance of a gold plate.</p> <p>&nbsp;</p> <p><strong>Dataset characterization</strong></p> <p>The database, named BSSL_DB_Key_Soils, has five sheets containing the key soil attributes, Vis-NIR-SWIR and&nbsp;MIR datasets, descriptions of the contributors and the proximal sensing methods used for spectral soil analysis. The sheets can be linked by &quot;ID_Unique&quot; columns, which bring the corresponding rows according to the data type. Some cells are empty&nbsp;because collaborators have already provided data in this way. However, we have decided to keep them&nbsp;in the database because they have other soil key attributes.&nbsp;Every Column in the data sheets is described as follows:</p> <p>&nbsp;</p> <p><strong>Sheet 1.&nbsp; &nbsp; &nbsp; &nbsp;BSSL_Soil_Attributes_Dataset</strong></p> <p>Column 1.&nbsp;&nbsp;&nbsp;&nbsp;<strong>ID_unique</strong>: Sequential code assigned to every record;</p> <p>Column 2.&nbsp;&nbsp;&nbsp;&nbsp;<strong>Owner code</strong>: Acronym assigned to each contributor who allowed access to their proprietary data;</p> <p>Column 3.&nbsp;&nbsp;&nbsp;&nbsp;<strong>Vis_NIR_SWIR_availability</strong>: availability of spectral data in visible, near-infrared, and shortwave infrared ranges;</p> <p>Column 4.&nbsp;&nbsp;&nbsp;&nbsp;<strong>MIR_availability</strong>: availability of spectral data in the middle infrared range;</p> <p>Column 5.&nbsp;&nbsp;&nbsp;&nbsp;<strong>Sampling</strong>: type of soil sampling;</p> <p>Column 6.&nbsp;&nbsp;&nbsp;&nbsp;<strong>Depth_cm</strong>: soil surface layer depth in centimeters;&nbsp;&nbsp;</p> <p>Column 7.&nbsp; &nbsp;&nbsp;<strong>Lat</strong>: Latitude;&nbsp;&nbsp;</p> <p>Column 8.&nbsp; &nbsp;&nbsp;<strong>Lat</strong>: Longitude;&nbsp;&nbsp;</p> <p>Column 9.&nbsp;&nbsp;&nbsp;&nbsp;<strong>Region</strong>: Brazilian geographical region of samples&#39; source;</p> <p>Column 10.&nbsp;&nbsp;&nbsp;&nbsp;<strong>Municipality</strong>: Brazilian municipality of samples&#39; source;</p> <p>Column 11.&nbsp;&nbsp;&nbsp;<strong>State</strong>: Brazilian Federation Unit of samples&#39; source;</p> <p>Column 12.&nbsp;<strong>Vegetation</strong>: type of vegetal covering;</p> <p>Column 13.&nbsp;<strong>Biome</strong>: groupings of ecosystems that share similar characteristics and span different regions;</p> <p>Column 14.&nbsp;<strong>Geology</strong>: type of rock matter from local soil sampling;</p> <p>Column 15.&nbsp;<strong>Sand_gkg</strong>: Content of the soil fraction with grain size between 2 and 0.053 mm, expressed in grams per kilogram;</p> <p>Column 16.&nbsp;<strong>Clay_gkg</strong>: Content of soil fraction with grain size smaller than 0.002 mm, expressed in grams per kilogram;</p> <p>Column 17.&nbsp;<strong>SOM_gkg</strong>: Soil organic matter content, expressed in grams per kilogram;</p> <p>Column 18.&nbsp;<strong>pH_H2O</strong>: Soil hydrogen ion potential measured in water;</p> <p>Column 19.&nbsp;<strong>Ca_mmolkg</strong>: Exchangeable calcium content in the soil, expressed in millimoles per kilogram;</p> <p>Column 20.&nbsp;<strong>Mg_mmolkg</strong>: Exchangeable magnesium content in the soil, expressed in millimoles per kilogram;</p> <p>Column 21.&nbsp;<strong>Na_mmolkg</strong>: Exchangeable sodium content in the soil, expressed in millimoles per kilogram; and</p> <p>Column 22.&nbsp;<strong>CEC_Ph7_mmolkg</strong>: Cation exchange capacity of the soil at neutral pH, expressed in millimoles per kilogram.</p> <p>&nbsp;</p> <p><strong>Sheet 2.&nbsp; &nbsp; &nbsp;BSSL_Vis_NIR_SWIR_Dataset</strong></p> <p>Column 1. <strong>ID_Unique</strong>: Sequential code assigned to every record;</p> <p>Column 2. <strong>Owner code</strong>: Acronym assigned to each contributor who allowed access to their proprietary data; and</p> <p>Column 3 &ndash; 2153. <strong>350 &ndash; 2500</strong>: Reflectance in 2151 spectral bands in nanometers from visible and near-infrared to shortwave infrared range (350 &ndash; 2500 nm).</p> <p>&nbsp;</p> <p><strong>Sheet 3.&nbsp; &nbsp;&nbsp;BSSL_MIR_Dataset</strong></p> <p>Column 1.&nbsp;<strong>ID_Unique:</strong>&nbsp;Sequential code assigned to every record;</p> <p>Column 2.&nbsp;<strong>Owner_code:</strong>&nbsp;Acronym assigned to each contributor who allowed access to their proprietary data; and</p> <p>Column 3 &ndash; 683.&nbsp;<strong>4000&nbsp;&ndash; 600:</strong>&nbsp;Reflectance in 681 spectral bands in centimeters in the middle infrared range (4000&nbsp;&ndash; 600 cm<sup>-1</sup>).</p> <p>&nbsp;</p> <p><strong>Sheet 4. &nbsp; &nbsp;&nbsp;&nbsp; Contributors</strong></p> <p>Column 1.&nbsp;<strong>Owner_code</strong>: Acronym assigned to each contributor who allowed access to their proprietary data, which identifies and links it to datasets;</p> <p>Column 2.&nbsp;<strong>Owner</strong>: Name of the collaborator who agreed to the availability of the data;</p> <p>Column 3.&nbsp;<strong>E-mail</strong>: Contact the e-mail of the owner for more information or a data request;</p> <p>Column 4.&nbsp;<strong>Institution</strong>: Contributor&#39;s affiliation;</p> <p>Column 5.&nbsp;<strong>Samples NIR</strong>: Number of Vis-NIR-SWIR samples sent to the BSSL collection;</p> <p>Column 6.&nbsp;<strong>Samples MIR</strong>: Number of MIR samples sent to the BSSL collection;</p> <p>&nbsp;</p> <p><strong>Sheet 5. &nbsp; &nbsp;&nbsp;&nbsp; Metadata</strong></p> <p>Column 1.&nbsp;<strong>Material and Methods</strong>: Description of procedures performed for soil data analyses</p> <p>&nbsp;</p> <p><strong>Expectation and Social Relevance</strong></p> <p>These data can impact various disciplines such as soil surveying, soil attribute mapping, soil analysis, soil mineralogy, soil management zones, precision agriculture, development of new datasets and scientific groups, and others. We expect this contribution to be valuable and useful to the soil research community in promoting this non-renewable natural resource&#39;s conservation and sustainable use.</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

A pangenome-guided manually curated library of transposable elements for Zymoseptoria tritici

<p>A manually-curated TE consensus library generated using a panel of 19 reference genomes for&nbsp;<em>Zymoseptoria tritici</em><sup>1-3</sup>&nbsp;along with reference genome assemblies for the sister species&nbsp;<em>Z. ardabiliae</em>,&nbsp;<em>Z. brevis</em>,&nbsp;<em>Z. pseudotritici</em>, and&nbsp;<em>Z. passerinii<sup>4</sup></em>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p>Putative TE consensus sequences were first obtained by annotating all 23 genome assemblies<sup>1&ndash;4</sup>&nbsp;with Earl Grey with default settings (v3.0;&nbsp;<a href="https://github.com/TobyBaril/EarlGrey">https://github.com/TobyBaril/EarlGrey</a>)<sup>5,6</sup>. Consensus sequences generated from each reference genome were clustered using CD-Hit-Est (v4.8.1)<sup>7,8</sup>&nbsp;to group sequences with 90% similarity across 80% of the longer sequence length (<em>-n 8 -d 0 -aL 0.8 -c 0.90 -G 0 -g 1 -b 500 -r 1</em>)&nbsp;to reduce redundancy whilst preventing the collapsing of chimeric sequences. Consensus sequences &lt;100bp were removed, as these are unlikely to represent true TE sequences. Each consensus sequence was then subject to manual curation as described by Goubert et al. (2022)<sup>9</sup>. Briefly, genomic copies of each TE were obtained using a &ldquo;BLAST, Extract, Extend&rdquo; process to recover genomic copies from each of the 23 reference genome assemblies with 1,000 flanking bases at either end<sup>9,10</sup>. For families with &gt;100 BLASTN hits, the 25 longest hits were selected, along with 75 random hits. Multiple alignments were generated for each putative TE family using MAFFT (v7.505) with the --auto flag<sup>11</sup>. Columns composed of &gt;=80% gaps were removed with T-COFFEE (v13.45.0.4846264)<sup>12</sup>. Subsequently, all sequence alignments were manually curated to define TE boundaries and remove regions of low conservation and rare insertions. Following manual curation, new majority-rule consensus sequences were generated with EMBOSS (v6.6.0.0) cons<sup>13</sup>. TE-Aid (<a href="https://github.com/clemgoub/TE-Aid/">https://github.com/clemgoub/TE-Aid/</a>) was used to aid visual inspection and to identify diagnostic features for classification of extended consensus sequences. Following this, TIRs were recorded if present, and nhmmscan (HMMER v3.3.2)<sup>14</sup>&nbsp;was used to identify homology to known curated elements in Dfam (v3.7). Combining this information, each TE consensus sequence was manually classified using available information following the naming convention &lsquo;&gt;ZymTri_2023_family_[n]#[Classification]/[Family]&rsquo;. Consensus sequences classified with low confidence have a &lsquo;?&rsquo; added to the name, as well as the string &lsquo;_LowConf&rsquo;. To reduce redundancy in the final TE library, sequences were clustered to the family-level using the 80-80-80 rule implemented in CD-hit-est<sup>9,15 </sup>(<em>-d 0 -aS 0.8 -c 0.8 -G 0 -g 1 -b 500 -r 1</em>). The representative sequence for each cluster was manually selected to select the sequence with the highest classification confidence, also defined as the &lsquo;most intact consensus&rsquo;. Chimeric sequences erroneously clustered were manually separated to retain sequences for the chimeric TE and the individual elements that generated the chimer.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Badet, T., Oggenfuss, U., Abraham, L., McDonald, B. A. &amp; Croll, D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen Zymoseptoria tritici.&nbsp;<em>BMC Biol.</em>&nbsp;<strong>18</strong>, 12 (2020).</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Goodwin, S. B.&nbsp;<em>et al.</em>&nbsp;Finished genome of the fungal wheat pathogen Mycosphaerella graminicola reveals dispensome structure, chromosome plasticity, and stealth pathogenesis.&nbsp;<em>PLoS Genet.</em>&nbsp;<strong>7</strong>, e1002070 (2011).</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Plissonneau, C., Hartmann, F. E. &amp; Croll, D. Pangenome analyses of the wheat pathogen Zymoseptoria tritici reveal the structural basis of a highly plastic eukaryotic genome.&nbsp;<em>BMC Biol.</em>&nbsp;<strong>16</strong>, 5 (2018).</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Feurtey, A.&nbsp;<em>et al.</em>&nbsp;Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria.&nbsp;<em>BMC Genomics</em>&nbsp;<strong>21</strong>, 588 (2020).</p> <p>5.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Baril, T., Imrie, R. M. &amp; Hayward, A. Earl Grey: a fully automated user-friendly transposable element annotation and analysis pipeline. (2022) doi:10.21203/rs.3.rs-1812599/v1.</p> <p>6.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Baril, T., Galbraith, J. &amp; Hayward, A.&nbsp;<em>Earl Grey</em>. (Zenodo, 2023). doi:10.5281/ZENODO.8116025.</p> <p>7.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Li, W. &amp; Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>22</strong>, 1658&ndash;1659 (2006).</p> <p>8.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Fu, L., Niu, B., Zhu, Z., Wu, S. &amp; Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>28</strong>, 3150&ndash;3152 (2012).</p> <p>9.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Goubert, C.&nbsp;<em>et al.</em>&nbsp;A beginner&rsquo;s guide to manual curation of transposable elements.&nbsp;<em>Mob. DNA</em>&nbsp;<strong>13</strong>, 7 (2022).</p> <p>10.&nbsp;&nbsp;&nbsp;Camacho, C.&nbsp;<em>et al.</em>&nbsp;BLAST+: Architecture and applications.&nbsp;<em>BMC Bioinformatics</em>&nbsp;<strong>10</strong>, 1&ndash;9 (2009).</p> <p>11.&nbsp;&nbsp;&nbsp;Katoh, K. &amp; Standley, D. M. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability.&nbsp;<em>Mol. Biol. Evol.</em>&nbsp;<strong>30</strong>, 772&ndash;780 (2013).</p> <p>12.&nbsp;&nbsp;&nbsp;Notredame, C., Higgins, D. G. &amp; Heringa, J. T-coffee: a novel method for fast and accurate multiple sequence alignment.&nbsp;<em>J. Mol. Biol.</em>&nbsp;<strong>302</strong>, 205&ndash;217 (2000).</p> <p>13.&nbsp;&nbsp;&nbsp;Rice, P., Longden, L. &amp; Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite.&nbsp;<em>Trends Genet.</em>&nbsp;<strong>16</strong>, 276&ndash;277 (2000).</p> <p>14.&nbsp;&nbsp;&nbsp;Wheeler, T. J. &amp; Eddy, S. R. nhmmer: DNA homology search with profile HMMs.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>29</strong>, 2487&ndash;2489 (2013).</p> <p>15.&nbsp;&nbsp;&nbsp;Wicker, T.&nbsp;<em>et al.</em>&nbsp;A unified classification system for eukaryotic transposable elements.&nbsp;<em>Nat. Rev. Genet.</em>&nbsp;<strong>8</strong>, 973&ndash;982 (2007).</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Optimal neutron-star mass ranges to constrain the equation of state of nuclear matter with electromagnetic and gravitational-wave observations: EOS library

<p>This repository includes a&nbsp;library of equations of state&nbsp;(EOS) and stellar models presented in the publications Weih et al. (2019) (see also the related identifier) and Most et al. (2018). The library&nbsp;includes ~ 3&nbsp;Million physically plausible EOSs that fulfill a number of astrophysical and nuclear constraints. See the README for more information.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Spectral library of laser-induced fluorescence (LiF) properties from Smithsonian rare-earth element (REE) orthophosphate standards

<p>The spectral library presents a data set of laser-induced fluorescence (LiF) spectra from rare-earth element (REE) orthophosphates provided and distributed as reference material for microbeam analysis by the Smithsonian National Museum of Natural History (sample IDs: 16484 - NMNH 168499; Jarosewich and Boatner, 1991; Donovan et al., 2002 and 2003). The data set delivers high-resolution LiF spectra excited at three standard laser wavelengths (325 nm, 442 nm, 532 nm) recorded in the UV-visible to near-infrared spectral range (340 - 1080nm). Presented LiF spectra represent data from efficient signal excitation conditions and contain the diagnostic emission lines of individual REE including detailed information on splitting into sub-levels. The LiF spectral library data provides a reference for various applications in spectroscopy-based material composition analysis with the scope of REE identification. LiF as a tool can complement the merging technique of reflectance spectroscopy, because LiF is a particularly well suited method for REE detection and can be used to cross-validate results (e.g. Lorenz et al. 2019) The LiF library allows for transparent and reproducible result analysis in scientific studies and promotes further developments of efficient automated algorithms for REE identification and characterisation. This addresses especially the need for innovative, non-invasive techniques of raw material exploration (securing REE supply) and material stream characterisation (e.g. in e-waste recycling) or for manifold applications in other fields of geosciences (e.g. geology) and physics.</p> <p>references:</p> <p>Donovan, J., Hanchar, J., Picolli, P., Schrier, M., Boatner, L., Jarosewich, E., 2002. Contamination in the rare-earth element orthophosphate reference sam- ples. J. Res. National Institute of Standards and Technology 106, 693&ndash;701. doi:10.6028/jres.107.056.&nbsp;</p> <p>Donovan, J., Hanchar, J., Piccoli, P., Schrier, M., Boatner, L., Jarosewich, E., 2003. A reexamination of the rare-earth element orthophosphate reference samples for electron microprobe analysis. Canadian Mineralogist 41, 221&ndash; 232. doi:10.2113/gscanmin.41.1.221.&nbsp;</p> <p>Jarosewich, E., Boatner, L., 1991. Rare-earth element reference samples for electron microprobe analysis. Geostandards Newsletter 15, 397&ndash;399. doi:10. 1111/j.1751-908X.1991.tb00115.x.&nbsp;</p> <p>Lorenz, S., Beyer, J., Fuchs, M., Seidel, P., Turner, D., Heitmann, J., Gloaguen, R., 2019. The Potential of Reflectance and Laser Induced Luminescence Spectroscopy for Near-Field Rare Earth Element Detection in Mineral Ex- ploration. Remote Sensing 11, 21. doi:10.3390/rs11010021.</p>

opencc-by-4.0Sep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record