Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

44

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

44 results for “GEMINI”

Learn how ShareScore rates datasets ↗
zenodo48/100

Programming Problems Submitted for Evaluation of LLMs GPT-3.5 and Gemini Pro 1.0

<p>Problems extracted from platforms LeetCode and BeeCrowd for evaluation of LLMs GPT3.5 and Gemini Pro 1.0.</p> <p>The data from the plataforms has the following columns and values:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <table> <tbody> <tr> <td><strong>LeetCode Data</strong></td> <td>&nbsp;</td> <td><strong>BeeCrowd Data</strong></td> <td>&nbsp;</td> </tr> <tr> <td><strong>Column</strong></td> <td><strong>Doc</strong></td> <td><strong>Column</strong></td> <td><strong>Doc</strong></td> </tr> <tr> <td>problem_level</td> <td>easy | medium | hard</td> <td>problem_level</td> <td>&lt;1...10&gt;</td> </tr> <tr> <td>problem_link</td> <td>&lt;LeetCode link for the problem&gt;</td> <td>problem_link</td> <td>&lt;BeeCrowd link for the problem&gt;</td> </tr> <tr> <td>prompt</td> <td>&lt;text submitted to LLM&gt;</td> <td>prompt</td> <td>&lt;text submitted to LLM&gt;</td> </tr> <tr> <td>response_code</td> <td>&lt;code provided by the LLM&gt;</td> <td>response_code</td> <td>&lt;code provided by the LLM&gt;</td> </tr> <tr> <td>response_evaluation</td> <td>True | False</td> <td>response_evaluation</td> <td>True | False</td> </tr> <tr> <td>execution_time_ms</td> <td>&lt;time&gt;</td> <td>execution_time_ms</td> <td>&lt;time&gt;</td> </tr> <tr> <td>memory_usage_mb</td> <td>&lt;memory&gt;</td> <td>error_generated</td> <td>Wrong Answer | Time Limit Exceeded | Memory Limit Exceeded....</td> </tr> <tr> <td>error_generated</td> <td>Wrong Answer | Time Limit Exceeded | Memory Limit Exceeded....</td> <td>attempts_number</td> <td>1 | 2 | 3</td> </tr> <tr> <td>attempts_number</td> <td>1 | 2 | 3</td> <td>author</td> <td>&lt;author's name&gt;</td> </tr> <tr> <td>contains_image</td> <td>True | False</td> <td>source</td> <td>&lt;origin institution&gt;</td> </tr> <tr> <td>related_topic_1</td> <td>&lt;topic&gt;</td> <td>origin_country</td> <td>&lt;origin country&gt;</td> </tr> <tr> <td>related_topic_2</td> <td>&lt;topic&gt;</td> <td>contains_image</td> <td>True | False</td> </tr> <tr> <td>related_topic_3</td> <td>&lt;topic&gt;</td> <td>related_topic_1</td> <td>&lt;topic&gt;</td> </tr> <tr> <td>related_topic_4</td> <td>&lt;topic&gt;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>related_topic_5</td> <td>&lt;topic&gt;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Associated data for "The Roasting Marshmallows Program with IGRINS on Gemini South II -- WASP-121 b has super-stellar C/O and refractory-to-volatile ratios" Published in The Astronomical Journal

<table> <tbody> <tr> <td>File Name</td> <td>Description</td> </tr> <tr> <td>w121_1DRC_H2O_ONLY.txt</td> <td>Self consistent, solar composition model spectrum with only H2O opacity.</td> </tr> <tr> <td>w121_1DRC_OH_ONLY.txt</td> <td>Self consistent, solar composition model spectrum with only OH opacity.</td> </tr> <tr> <td>w121_1DRC_CO_ONLY.txt</td> <td>Self consistent, solar composition model spectrum with only CO opacity.</td> </tr> <tr> <td>w121_1DRC_EVERYTHING.txt</td> <td>Self consistent, solar composition model spectrum with all sources of opacity.</td> </tr> <tr> <td>pre_.pic</td> <td>Pre-eclipse data in data cuboid of shape N_order, N_frame, N_pixel</td> </tr> <tr> <td>pre_variance.pic</td> <td>Associated per-pixel variance for the pre-eclipse data.</td> </tr> <tr> <td>post_cube.pic</td> <td>Post-eclipse data in data cuboid of shape N_order, N_frame, N_pixel</td> </tr> <tr> <td>pre_time_BJD.pic</td> <td>Average frame time in BJD for the pre-eclipse sequence.</td> </tr> <tr> <td>pre_rvel.pic</td> <td>Stellar radial velocity, including barycentric correction, per frame for the pre-eclipse sequence.</td> </tr> <tr> <td>post_ph.pic</td> <td>Orbital phase per frame for the post-eclipse sequence.</td> </tr> <tr> <td>post_time_BJD.pic</td> <td>Average frame time in BJD for the post-eclipse sequence.</td> </tr> <tr> <td>post_rvel.pic</td> <td>Stellar radial velocity, including barycentric correction, per frame for the post-eclipse sequence</td> </tr> </tbody> </table>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Dados abertos do artigo 'Inteligência artificial no levantamento bibliográfico em bases de dados científicos: comparando expressões de busca no ChatGPT, Copilot e Gemini'

<p>Resultado por IA de todas os comandos executados na pesquisa. Texto em formado PDF.<br><br><br>Artigo dispon&iacute;vel em &rarr; https://doi.org/10.20396/rdbci.v23i00.8678378 ou https://periodicos.sbu.unicamp.br/ojs/index.php/rdbci/article/view/8678378</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Gemini meteor shower, by Hao Yin, China

<p>Third place in the 2021 IAU OAE Astrophotography Contest, category Meteor showers.</p> <p>As the Earth travels around the Sun, it may cross the path of debris left behind by a comet or, more rarely, by an asteroid. These debris enter the atmosphere at high speed, producing beautiful tracks as they burn in the sky due to friction with the atmosphere. The image captures the Geminid meteor shower, named because the radiant point is located on the sky in the constellation Gemini. The particles composing the meteor shower travel at similar speed and in parallel trajectories, which causes a perspective effect like if the stream radiates from one single point in the sky, which is known as the radiant point. This image, taken in December 2020 in China, clearly shows this perspective. This is a very prolific shower, in such a way that over one hundred meteorites could be seen per hour in recent appearances. This meteor shower is one of the few associated not with a comet, but with an asteroid &ndash; 3200 Phaeton, which might be a comet that lost all its volatile material. This image shows the large number of meteors that can be observed in this shower, which always happens in December every year. The image also shows one of the most prominent constellations in the night sky, Orion, easily seen by the three stars in a diagonal making up Orion&rsquo;s Belt, and the red-orange star Betelgeuse. Right above the dish is a bright point and that is the Sirius, the brightest star in the night sky and part of the constellation Canis Major. The fuzzy bluish smudge at around 2 &lsquo;o&rsquo; clock is the Pleiades star cluster.</p> <p>Credit:&nbsp;Hao Yin/IAU OAE</p>

opencc-by-4.0Aug 2021View details →
ClinicalTrials.gov40/100

Relapsing Forms of Multiple Sclerosis (RMS) Study of Bruton's Tyrosine Kinase (BTK) Inhibitor Tolebrutinib (SAR442168) (GEMINI 1)

ClinicalTrials.gov study NCT04410978. IPD Sharing: YES. Countries: 24. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov40/100

Relapsing Forms of Multiple Sclerosis (RMS) Study of Bruton's Tyrosine Kinase (BTK) Inhibitor Tolebrutinib (SAR442168) (GEMINI 2)

ClinicalTrials.gov study NCT04410991. IPD Sharing: YES. Countries: 28. Publications: 1.

controlledIPD-YESFeb 2026View details →
zenodo36/100

Code on Demand: A Comparative Analysis of the Efficiency, Understandability, and Self-Correction Capability of Copilot, ChatGPT, and Gemini - Data resulting from the study

<p>Este conjunto de dados foi gerado como parte do estudo "Code on Demand: A Comparative Analysis of the Efficiency, Understandability, and Self-Correction Capability of Copilot, ChatGPT, and Gemini - Data resulting from the study". O estudo focou na avalia&ccedil;&atilde;o do desempenho das ferramentas Copilot, ChatGPT e Gemini, utilizando problemas do LeetCode em quatro linguagens de programa&ccedil;&atilde;o: Python, Java, JavaScript e C.</p> <p>O conjunto de dados atualizado est&aacute; organizado nas seguintes pastas:</p> <ol> <li> <p><strong>c_programs</strong>: Esta pasta cont&eacute;m os scripts Python utilizados para calcular a complexidade ciclom&aacute;tica e a complexidade cognitiva do c&oacute;digo C gerado pelas ferramentas.</p> <ul> <li><code>calculate_cyclomatic_complexity.py</code>: Script para calcular a complexidade ciclom&aacute;tica.</li> <li><code>calculate_cognitive_complexity.py</code>: Script para calcular a complexidade cognitiva.</li> </ul> </li> <li> <p><strong>codes_suggested_by_the_tools</strong>: Esta pasta cont&eacute;m as sugest&otilde;es de c&oacute;digo geradas pelo Copilot, ChatGPT e Gemini para cada problema do LeetCode.</p> <ul> <li>Subpastas: <code>ChatGPT</code>, <code>Copilot</code>, <code>Gemini</code>, cada uma contendo as sugest&otilde;es de c&oacute;digo correspondentes nos formatos das linguagens.</li> </ul> </li> <li> <p><strong>complexity_of_codes</strong>: Esta pasta cont&eacute;m dois arquivos CSV que fornecem os resultados da an&aacute;lise de complexidade para o c&oacute;digo gerado.</p> <ul> <li><code>AI analysis results table - Cognitive.csv</code>: Resultados da complexidade cognitiva do c&oacute;digo gerado.</li> <li><code>AI analysis results table - Cyclomatic.csv</code>: Resultados da complexidade ciclom&aacute;tica do c&oacute;digo gerado.</li> </ul> </li> </ol> <p>Este conjunto de dados atualizado oferece insights valiosos sobre o desempenho das ferramentas de gera&ccedil;&atilde;o de c&oacute;digo com IA e pode ser utilizado para an&aacute;lises futuras ou estudos de replica&ccedil;&atilde;o.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

GEMINI onco-extras package

<p>Oncology-specific annotation files for use with GEMINI - with (upcoming) special support through Galaxy&#39;s GEMINI annotate tool wrapper.</p>

openodc-odblMar 2019View details →
ClinicalTrials.gov36/100

An Efficacy, Safety, and Tolerability Study Comparing Dolutegravir (DTG) Plus Lamivudine (3TC) With Dolutegravir Plus Tenofovir/Emtricitabine in Treatment naïve HIV Infected Participants (Gemini 2)

ClinicalTrials.gov study NCT02831764. IPD Sharing: YES. Countries: 18. Publications: 3.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

An Efficacy, Safety, and Tolerability Study Comparing Dolutegravir Plus Lamivudine With Dolutegravir Plus Tenofovir/Emtricitabine in Treatment naïve HIV Infected Subjects (Gemini 1)

ClinicalTrials.gov study NCT02831673. IPD Sharing: Not stated. Countries: 18. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Genomic Medicine for Ill Neonates and Infants (The GEMINI Study)

ClinicalTrials.gov study NCT03890679. IPD Sharing: YES. Countries: 1. Publications: 3.

controlledIPD-YESFeb 2026View details →
zenodo32/100

Gemini 3D test data

<p>DEPRECATED: data now at:&nbsp;<a href="https://zenodo.org/record/3962801">https://zenodo.org/record/3962801</a></p> <p>---</p> <p>Gemini 3D simulation reference data</p> <p>https://www.github.com/gemini3d/gemini</p> <p>v3.2.0 new api / meta</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Multifrequency study of HH 137 and HH 138: Discovering new knots and molecular outflows with Gemini and APEX

<p>We present reduced H<sub>2</sub> and K near-infrared band images of HH 137, taken with the GSAOI+GeMS&nbsp;instrument of the Gemini South telescope. In addition, we provide submilimiter data in <sup>12</sup>CO(3-2),&nbsp;<sup>13</sup>CO(3-2), C<sup>18</sup>O(3-2), HCO<sup>+</sup>(3-2) and HCN(3-2) molecular lines of both HH 137 and HH 138,&nbsp;obtained with the SHeFI instrument of the Atacama Pathfinder EXperiment (APEX) telescope.</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Supplemental data: Benchmarking of Exomiser and seven LLMs (GPT o1 preview, GPT o1 mini, GPT-4o, Gemini Flash 2.0, Meditron-70B, Meditron3-70B, and Medfound-175B) for differential diagnosis using phenopackets

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Gemini 1.0 Pro, Claude 3 Sonnet, Microsoft Copilot, and ChatGPT-4 responses on the Test of Understanding Graphs in Kinematics (TUG-K), April 2024

<p>The data contains 30 responses from each chatbot to 26 items on the TUG-K survey. The chatbots tested were Google Gemini (freely available version, Gemini Pro 1.0), Claude 3 Sonnet, Microsoft Copilot (freely available version, balanced setting) and ChatGPT-4 (subscription-based, ChatGPT Plus). The prompts consisted of screenshots of the test items and the sentence "Answer the question in the image" for Copilot and Gemini 1.0 Pro. For Claude 3 Sonnet and ChatGPT-4 the prompt consisted of the screenshot only. The data was collected in April 2024.</p> <p>The data is a continuation of the research data on ChatGPT-4's performance on the TUG-K (10.5281/zenodo.10429075).</p>

opencc-by-4.0May 2024View details →
zenodo32/100

ChatGPT-4o, Claude 3 Opus, Gemini 1.0 Ultra, Gemini 1.5 Pro, and ChatGPT-4 responses on the Test of Understanding Graphs in Kinematics (TUG-K), April 2024

<p>The chatbots tested were Google's Gemini 1.0 Ultra, Google's Gemini 1.5 Pro (prompted through Google AI studio using default settings), Anthropic's Claude 3 Opus, OpenAI's ChatGPT-4o (using the latest model GPT-4o) and OpenAI's ChatGPT-4 (using the GPT-4 model). The prompts consisted only of screenshots of the test items. For Gemini 1.0 Ultra, the image was accompanied by the sentence "Answer the question in the image".</p> <p>The data, collected in April 2024, contains 30 responses from each chatbot to 26 items on the TUG-K survey. For ChatGPT using the GPT-4 model (ChatGPT-4), we provide two separate datasets. One complete dataset (all TUG-K items) without the use "advanced data analysis" plugin, and one with only those six items where the "advanced data analysis" plugin was automaticaly used by the chatbot. These two datasets partially overlap with the ChatGPT-4 dataset published previously (see link below).</p> <p>This dataset focuses on subscription-based chatbots and is a continuation of a previous dataset that focused on freely available chatbots(10.5281/zenodo.11183803).</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Gemini North GMOS Dataset

<p>Gemini GMOS observation of U20171.</p>

opencc-by-4.0Nov 2018View details →
zenodo32/100

Replication package of the paper "Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot"

<h1>Replication Package</h1> <p>This replication package contains the necessary tools, data, and scripts for reproducing the results of our paper: "<em>Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot</em>". Below is a detailed description of the directory structure and the contents of this package.</p> <h2>Contents</h2> <p>The replication package is organized into two main directories:</p> <ul> <li> <p><code>assets</code>: This directory contains all .csv files used as input for the script and the outputted .csv file used to perform the manual and automated analyses for RQ1 and RQ2.</p> </li> <li> <p><code>script</code>: This directory contains all scripts for RQ1 and RQ2.</p> </li> </ul> <p>In the following, we describe the content of each directory:</p> <h2><code>assets</code></h2> <p>This directory contains the tools and resources required for our study.</p> <h3><code>dataset</code>: Contains the main datasets used in the study.</h3> <ul> <li> <p><code>annotationStore.csv</code>: Input dataset for our analyses, originating from the <em>CODESEARCHNET</em> dataset.</p> </li> <li> <p><code>queries.csv</code>: .csv file containing the queries used for the experiments filtered from the <em>CODESEARCHNET</em>dataset. This file contains the following columns:</p> <ul> <li><em>Language</em>: Programming language of the query</li> <li><em>Query</em>: Query used for the experiment</li> <li><em>GitHubUrl</em>: GitHub URL related to a snippet that addresses the query</li> <li><em>Relevance</em>: Relevance of the linked GitHub snippet to the query</li> </ul> </li> </ul> <h3><code>data</code>: Contains the datasets and results of all analyses.</h3> <ul> <li> <p><code>queries.csv</code>: General input queries. This file contains the following columns:</p> <ul> <li><em>Language</em>: Programming language of the query</li> <li><em>Query</em>: Query used for the snippet generation</li> <li><em>Prompt</em>: LLM prompt generated for the query as: <em>You are a Senior <code>&lt;Language&gt;</code> developer. Then give me a <code>&lt;Language&gt;</code> code snippet about: <code>&lt;Query&gt;</code></em></li> </ul> </li> <li> <p><code>queries_filled.csv</code>: Similar to the previous file, but also containing the output produced by the LLM-based assistants. This file contains the following columns:</p> <ul> <li><em>Language</em>: Programming language of the query</li> <li><em>Query</em>: Query used for the snippet generation</li> <li><em>Prompt</em>: LLM prompt generated for the query as: <em>You are a Senior <code>&lt;Language&gt;</code> developer. Then give me a <code>&lt;Language&gt;</code> code snippet about: <code>&lt;Query&gt;</code></em></li> <li><em>Notes</em>: General notes that provide additional context or information about the query or prompt.</li> <li><em>Gemini_Answer(n)</em>: The generated code snippets by Gemini.</li> <li><em>Gemini(n)</em>: The external links provided by Gemini.</li> <li><em>Prompt (repeated)</em></li> <li><em>Note</em>: Notes that provide additional context or information about the query or prompt.</li> <li><em>Copilot_Answer(n)</em>: The generated code snippets by Bing-Copilot.</li> <li><em>Copilot_Bing(n)</em>: The external links provided by Bing-Copilot.</li> </ul> </li> </ul> <h4><code>copilot</code> || <code>gemini</code>: Contains the data related to the specific LLM. These two subdirectories have the same internal structure.</h4> <ul> <li><code>queries.csv</code>: The <code>queries_filled.csv</code> file, filtered for the specific LLM.</li> <li><code>queries_noTrivial.csv</code>: Contains only the queries with at least one nontrivial generated snippet.</li> <li> <p><code>external_links.csv</code>: External links extracted from the LLMs output.</p> </li> <li> <p><code>external_links_filled.csv</code>: Snippets extracted from the external links.</p> <ul> <li><em>index</em>: Query ID</li> <li><em>source</em>: Snippet ID</li> <li><em>url</em>: Link URL</li> <li><em>note</em>: Notes that provide additional context or information about the query or prompt</li> <li><em>code(n)</em>: The n-th code snippet extracted from the source</li> </ul> </li> </ul> <h4><code>manual_analysis</code>: Manual analysis results.</h4> <ul> <li><code>manual_analysis.csv</code>: <ul> <li><em>index</em>: Query ID</li> <li><em>query</em>: Query used for the snippet generation</li> <li><em>generatedsnippet(n)</em>: The n-th code snippet generated by the LLM-based assistant</li> <li><em>trivial_1</em>: Manual analysis of whether or not the snippet was trivial (validator 1)</li> <li><em>trivial_2</em>: Manual analysis of whether or not the snippet was trivial (validator 2)</li> <li><em>trivial_final</em>: Manual analysis of whether or not the snippet was trivial (final classification if there is a disagreement)</li> <li><em>source</em>: URL to analyze</li> <li><em>sourcetype1</em>: Type of the source (validator 1)</li> <li><em>sourcetype2</em>: Type of the source (validator 2)</li> <li><em>sourcetypefinal</em>: Type of the source (final classification if there is a disagreement)</li> <li><em>relatedtoquery_1</em>: Relevance of the link to the query (validator 1)</li> <li><em>relatedtoquery_2</em>: Relevance of the link to the query (validator 2)</li> <li><em>relatedtoquery_final</em>: Relevance of the link to the query (final classification if there is a disagreement)</li> <li><em>relatedtosnippets_1</em>: Relevance of the generated snippet to those in the link (validator 1)</li> <li><em>relatedtosnippets_2</em>: Relevance of the generated snippet to those in the link (validator 2)</li> <li><em>relatedtosnippets_final</em>: Relevance of the generated snippet to those in the link (final classification if there is a disagreement)</li> </ul> </li> <li><code>manual_analysis_noTrivial.csv</code>: As in the previous file, but only the queries with at least one nontrivial generated code snippet.</li> </ul> <h4><code>clone_detector</code>: Output and intermediate files for clone detection with Copilot data.</h4> <ul> <li><code>copilot_tokens || gemini_tokens</code>: Contains the output the tokenization of the generated code snippets and the code snippets extracted from the external links.</li> <li><code>merged_llm_ext_link.csv</code>: All possible pairs (Cartesian product) (code snippet extracted from the external links, generated code snippet). This file is the input of the clone detection tool. <ul> <li><em>ID_query</em>: Query ID</li> <li><em>query</em>: Query used for the snippet generation</li> <li><em>language</em>: Programming language of the query</li> <li><em>generated_snippet</em>: The generated code snippet by the LLM-based assistant</li> <li><em>IDgensnippet</em>: The index of the generated code snippet</li> <li><em>LOCgensnippet</em>: The number of lines of code of the generated code snippet</li> <li><em>ID_source</em>: Source ID</li> <li><em>source</em>: Source URL</li> <li><em>source_snippet</em>: Code snippet extracted from the source</li> <li><em>IDsourcesnippet</em>: ID of the code snippet extracted from the source</li> <li><em>LOCsourcesnippet</em>: The number of lines of code of the code snippet extracted from the source</li> <li><em>note</em>: Notes that provide additional context or information about the query or prompt</li> </ul> </li> <li><code>clone_detection_output.csv</code>: Contains the clone detection results. <ul> <li><em>ID_query</em>: The index of the query</li> <li><em>query</em>: Query used for the snippet generation</li> <li><em>language</em>: The programming language of the query</li> <li><em>generated_snippet</em>: The generated code snippet by the LLM-based assistant</li> <li><em>IDgensnippet</em>: The index of the generated code snippet</li> <li><em>LOCgensnippet</em>: The number of lines of code of the generated code snippet</li> <li><em>ID_source</em>: Source ID</li> <li><em>source</em>: Source URL</li> <li><em>source_snippet</em>: Code snippet extracted from the source</li> <li><em>IDsourcesnippet</em>: ID of the code snippet extracted from the source</li> <li><em>LOCsourcesnippet</em>: The number of lines of code of the code snippet extracted from the source</li> <li><em>note</em>: Notes that provide additional context or information about the query or prompt</li> <li><em>clone_detected</em>: bBolean value that indicates whether a clone has been detected (1 = detected, 0 = not detected)</li> <li><em>cloning_ratio</em>: Ratio of the number of lines of code of the generated code snippet has been detected as a clone in the code snippet extracted from the source</li> <li><em>cloned_lines</em>: The number of lines of code of the generated code snippet that has been detected as a clone in the code snippet extracted from the source</li> </ul> </li> </ul> <h4><code>cosine_sim</code>: Cosine similarity results.</h4> <ul> <li><code>cosine_sim_output.csv</code>: Contains the cosine similarity results <ul> <li><em>query_id</em>: Query ID</li> <li><em>snippet_id</em>:ID the generated code snippet</li> <li><em>source_id</em>: ID of the source</li> <li><em>sourcesnippetid</em>: ID of the code snippet extracted from the source <ul> <li><em>cosine_similarity</em>: The cosine similarity between the generated code snippet and the code snippet extracted from the source</li> </ul> </li> </ul> </li> </ul> <h4><code>quant_analysis</code>: Quantitative analysis results.</h4> <ul> <li><code>topN_links_se.csv</code>: Contains the top-N links extracted from the search engine. <ul> <li><em>id</em>: Query ID</li> <li><em>query</em>: The query</li> <li><em>url</em>: Link URL</li> </ul> </li> <li><code>merged_clone_cosine.csv</code>: Contains the merged results of the clone detection and cosine similarity. <ul> <li><em>ID_query</em>: Query ID</li> <li><em>query</em>: The query</li> <li><em>language</em>: The programming language of the query</li> <li><em>generated_snippet</em>: The generated code snippet by the LLM-based assistant</li> <li><em>IDgensnippet</em>: The ID of the generated code snippet</li> <li><em>LOCgensnippet</em>: The number of lines of code of the generated code snippet</li> <li><em>ID_source</em>: The index of the source</li> <li><em>source</em>: The source URL</li> <li><em>source_snippet</em>: The code snippet extracted from the source</li> <li><em>IDsourcesnippet</em>: The index of the code snippet extracted from the source</li> <li><em>LOCsourcesnippet</em>: The number of lines of code of the code snippet extracted from the source</li> <li><em>note</em>: Notes that provide additional context or information about the query or prompt</li> <li><em>clone_detected</em>: Boolean value that indicates if a clone has been detected(1 = detected, 0 = not detected)</li> <li><em>cloning_ratio</em>: The ratio of the number of lines of code of the generated code snippet has been detected as a clone in the code snippet extracted from the source</li> <li><em>cloned_lines</em>: The number of lines of code of the generated code snippet that has been detected as a clone in the code snippet extracted from the source</li> <li><em>cosine_similarity</em>: The cosine similarity between the generated code snippet and the code snippet extracted from the source</li> </ul> </li> </ul> <h4><code>other_analysis</code>: Contains more performed analysis.</h4> <ul> <li><code>sample_queries.csv</code>: Contains a sample of five queries for language used for perform the chain of thought experiment. <ul> <li><em>Language</em>: Programming language of the query</li> <li><em>Query</em>: Query used for the snippet generation</li> <li><em>Prompt</em>: LLM prompt generated for the query as: <em>You are a Senior <code>&lt;Language&gt;</code> developer. Then give me a <code>&lt;Language&gt;</code> code snippet about: <code>&lt;Query&gt;</code></em></li> </ul> </li> <li><code>chain_of_thought.csv</code>: <ul> <li><em>Language</em>: Programming language of the query</li> <li><em>Query</em>: Query used for the snippet generation</li> <li><em>Prompt</em>: LLM prompt generated for the query as: <em>You are a Senior <code>&lt;Language&gt;</code> developer. Then give me a <code>&lt;Language&gt;</code> code snippet about: <code>&lt;Query&gt;</code></em></li> <li><em>Clone_fonud</em>: Boolean value that indicates if a clone has been detected (Yes = detected, No = not detected)</li> <li><em>Note</em>: Notes that provide additional context or information about the performed analysis</li> </ul> </li> <li><code>data_check.csv</code>: <ul> <li><em>Link</em>: URL of the source provided by the LLM</li> <li><em>Post_date</em>: Indicates if the date of the post is before/after the date of training of the LLM (before 2023, after 2023, not provided)</li> <li><em>Note</em>: Notes that provide additional context or information about the performed analysis</li> </ul> </li> </ul> <h4><code>results</code>: Final analysis results.</h4> <ul> <li><code>jaccard_analysis.csv</code>: Contains the results of the Jaccard analysis comparing the provided external links by the LLMs with the top-N links extracted from the corresponding search engine. <ul> <li><em>id</em>: Query ID</li> <li><em>language</em>: The programming language of the query</li> <li><em>llm_link</em>: The external links provided by the LLM</li> <li><em>llmlinksize</em>: The number of external links provided by the LLM</li> <li><em>overlap_links</em>: The overlapping links between the LLM and the search engine</li> <li><em>overlap_size</em>: The number of overlapping links between the LLM and the search engine</li> <li><em>nonoverlaplinks</em>: The non-overlapping links between the LLM and the search engine</li> <li><em>union_size</em>: The size of the union set links between the LLM and the search engine</li> <li><em>jaccard</em>: The Jaccard similarity between the LLM and the search engine</li> </ul> </li> <li><code>merged_analysis.csv</code>: Contains the merged results of the manual and quantitative analyses. <ul> <li><em>id</em>: The index of the query</li> <li><em>query</em>: The query used for the experiment</li> <li><em>trivial_final(n)</em>: The final assignment for the triviality of the n-th generated code snippet</li> <li><em>source</em>: The URL of the source</li> <li><em>sourcetypefinal</em>: The final assignment for the type of the source</li> <li><em>relatedtoquery_final</em>: The final assignment for the relevance of the generated code snippet to the query</li> <li><em>relatedtosnippets_final</em>: The final assignment for the relevance of the generated code snippet to the source</li> <li><em>cloning_ratio</em>: The maximum cloning ratio between the generated code snippet and all the code snippets extracted from the source</li> <li><em>cosine_similarity</em>: The cosine similarity related to the snippets with maximum cloning ratio</li> </ul> </li> </ul> <h3><code>cccfindersw-configuration-files</code>:</h3> <p>Contains additional configuration files for the CCFinderSW clone detection tool. The files are <code>javascript_comment.txt</code> and <code>javascript_reserved.txt</code>. They must be placed in the tool's <code>comment/</code> and <code>reserved/</code> directories.</p> <h3><code>appendix.tex</code>: The appendix of the paper containing:</h3> <ul> <li><em>Table 1</em>: Number of links of different types provided by Gemini and Bing CoPilot</li> </ul> <h3><code>appendix.pdf</code>: The appendix of the paper in PDF format.</h3> <h2><code>script</code></h2> <p>This directory contains our scripts (mostly Python, an R script and an Applescript) to preprocess data and run the clone detection analyses.</p> <ul> <li><code>1_dateset_filtering.py</code>: Script to filter the dataset. The input of this script is the <code>annotationStore.csv</code> file, and the output is the <code>queries.csv</code> file.</li> <li><code>2_prompt_generation.py</code>: Script to generate prompts for the LLM-based assistants. The input of this script is the <code>queries.csv</code> file, and the output is the <code>queries_filled.csv</code> file.</li> <li><code>3_gen_sheet_sources_extraction.py</code>: Script to split the external links provided by the LLM, one for each row. The input of this script is the <code>queries_filled.csv</code> file, and the output is the <code>external_link.csv</code> file.</li> <li><code>4_ext_link_snippet_extraction.py</code>: Script to extract the snippets from Web URLs. It only works for the most popular domains. The input of this script is the <code>external_links.csv</code> file, and the output is the <code>external_links_filled.csv</code> file.</li> <li><code>5_top_n_link_SearchEngine.py</code>: Script to perform top-N link search using the corresponding search engines (Google Search and Bing). The input of this script is the <code>queries_filled.csv</code> file. It executes the <code>browser_bot.scpt</code>. The output is the <code>topN_links_se.csv</code> file. <ul> <li><code>browser_bot.scpt</code>: Script for browser automation (AppleScript).</li> </ul> </li> <li><code>6_se_vs_llm.py</code>: Script to compare (using the Jaccard metric) the links returned by the corresponding search engines with those provided by the LLM-based assistants. The input of this script is the <code>external_links.csv</code> file and the <code>topN_links_se.csv</code> file. The output is the <code>jaccaard_analysis.csv</code> file.</li> <li><code>7_results_manual_analysis.py</code>: Script to extract results and statistical analyses performed on the manual analysis and reported in the tables in the paper.</li> <li><code>8_merge_gen_source_snippets.py</code>: This script takes as input: <code>{llm}/queries.csv</code> and <code>{llm}/external_link_filled.csv</code> to merge them and generates an expanded one, i.e., one in which we have on each line a snippet extracted from the source, this will be the input of our final script for clone detection. The output is the <code>merged_llm_ext_link.csv</code> file.</li> <li><code>9_clone_detection.py</code>: Script to perform clone detection. The input of this script is the <code>merged_llm_ext_link.csv</code> file, and the output is the <code>clone_detection_output.csv</code> file.</li> <li><code>10_cosine_sim_check.py</code>: Script to compute the code snippets' cosine similarity. The script takes as input the tokenized files from the <code>{llm}_tokens</code> directory. The output is the <code>cosine_sim_output.csv</code>file.</li> <li><code>11_merger_clone_cosine.py</code>: Script to merge clone detection and cosine similarity results. The input of this script is the <code>clone_detection_output.csv</code> and the <code>cosine_sim_output.csv</code> files, and the output is the <code>merged_clone_cosine.csv</code> file.</li> <li><code>12_merge_manual_quantitative_analysis.py</code>: Script to merge manual and quantitative analysis results. The inputs of this script are the <code>manual_analysis.csv</code> and the <code>merged_clone_cosine.csv</code>files, and the output is the <code>merged_analysis.csv</code> file.</li> <li><code>13_sample_for_COT_analysis.py</code>: Script to collect the sample of queries on which we perform the chain of thought analysis, the output is the <code>sample_queries.csv</code> file.</li> <li><code>14_llm_vs_csn.py</code>: Script to check the overlap between the links provided by the LLM and the one associated with the related query in the <em>CodeSearchNet</em> dataset. The input of this script is the <code>queries.csv</code>and <code>queries_noTrivial.csv</code> file.</li> <li><code>cloningGraph.R</code>: R script to generate the cloning graph. The input of this script is the <code>merged_analysis.csv</code> file.</li> </ul>

opencc-by-4.0Nov 2024View details →
ClinicalTrials.gov32/100

Global BurdEn of MechanIcal VeNtilatIon (GEMINI). VeNtilatIon (GEMINI Study) 2022 for VENTILAGROUP.

ClinicalTrials.gov study NCT05392010. IPD Sharing: NO. Countries: 1. Publications: 17.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Our Whole Lives Gemini: Virtual Integrative Medical Group Visits for Managing Chronic Pain

ClinicalTrials.gov study NCT06515925. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record