Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

267

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

267 results for “Patentes”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data for: 3D bioprinting patents

<p>This dataset contains information regarding 3D bioprinting patent/patent applications.&nbsp;</p> <p><a href="https://www.orbit.com/">Orbit</a>&nbsp;(a fee-based patent database provided by Questel), accessed on Jan. 2, 2022, was used for data mining.</p> <p>The files titled &ldquo;<em>3D bioprinting patents</em>&rdquo; and &ldquo;<em>Bioink patents</em>&rdquo; contain information related to Priority, Application and Publication numbers, Priority Application and Publication dates, Title, Abstract, and Current assignees.</p> <p>The patent searches were carried out by keywords and classification codes.&nbsp;</p> <p>Both IPC (International Patent Classification) and CPC (Cooperative Patent Classification) codes were used.&nbsp;</p> <p>Instead of 309 patents (see reference 1), a total number of 3,681 documents were retrieved (of which 3,027 are still alive and 2,461 filed in the period 2016 &ndash; 2020).</p> <p>Most of the published patent applications were generated in China and the USA (1,360 vs 1,063 priority applications). China is effectively the leading country (37%), followed by the USA with 29% of priority patent applications.&nbsp;</p> <p><strong>Value of the dataset</strong>: prior art searches, technological trends&nbsp;</p> <p><strong>Steps to reproduce data</strong>:&nbsp;</p> <p>The search strategy is reported in the table below:</p> <table> <tbody> <tr> <td> <p>1</p> </td> <td> <p>857</p> </td> <td> <p>(BIOPRINT+ OR BIOINK? OR ORGAN_ON_A_CHIP)/TI/AB/CLMS/ICLM</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>3096</p> </td> <td> <p>((A61L-027+ OR A61F-002+ OR A61L2430/00)</p> <p>AND (B33Y+ OR B29C-064+))/IPC/CPC</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>3681</p> </td> <td> <p>&nbsp;&nbsp;1 OR&nbsp;&nbsp;&nbsp;2</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>71912</p> </td> <td> <p>(C09D-011+)/IPC/CPC</p> </td> </tr> <tr> <td> <p>5</p> </td> <td> <p>104</p> </td> <td> <p>&nbsp;&nbsp;3 AND&nbsp;&nbsp;&nbsp;4</p> </td> </tr> </tbody> </table> <p><strong>Definition of the classification codes used</strong></p> <p><strong>A61L 27</strong>: Materials for grafts or prostheses or for coating grafts or prostheses</p> <p><strong>A61F 2</strong>: Filters implantable into blood vessels; Prostheses, i.e., artificial substitutes or replacements for parts of the body; Appliances for connecting them with the body; Devices providing patency to, or preventing collapsing of, tubular structures of the body, e.g., stents</p> <p><strong>A61L2430/00</strong>: Materials or treatment for tissue regeneration</p> <p><strong>B33Y</strong>: Additive manufacturing, i.e., manufacturing of three-dimensional [3-d] objects by additive deposition, additive agglomeration, or additive layering, e.g., by 3-d printing, stereolithography, or selective laser sintering</p> <p><strong>B29C 64/00</strong>: Additive manufacturing, i.e., manufacturing of three-dimensional [3D] objects by additive deposition, additive agglomeration, or additive layering, e.g., by 3D printing, stereolithography, or selective laser sintering</p> <p><strong>C09D 11</strong>: inks&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

OntoChem PFAS CORE and Patent Files for MetFrag

<p>These are MetFraggable versions of the PFAS lists produced by OntoChem by performing literature mining of the CORE database (27K entries) and Google Patent collections (1.7M entries). The MetFrag versions have undergone filtering to remove entries that prevent MetFrag running (i.e. multiple-entry formulas and certain elements). Each file contains a tag whether the PFAS fits three definitions, A, B and C. The CORE database also contains the number of references in which the PFAS entry was found. Each file is available as CSV or in compressed form.</p> <ul> <li>PFAS Definition A: Each compound that contains a CF<sub>2</sub> group</li> <li>PFAS Definition B: Each compound that contains a (AH)(AH)(F)C-C(AH)F<sub>2</sub> group, where AH groups could be hydrogen or any other atom and the bond between both aliphatic carbon atoms is a single bond</li> <li>PFAS Definition C: Each compound that contains a (R<sup>1</sup>)(R<sup>2</sup>)(F)C-C(R<sup>3</sup>)F<sub>2</sub> group is considered a PFAS, where the R groups are any atom except hydrogen and the bond between both aliphatic carbon atoms is a single bond</li> </ul> <p><strong><em>Please note these files are very large (especially patents) and should not be uploaded to MetFragWeb directly - they will be available from the dropdown menu. These are provided for command line users, and any other workflows interested in these files!&nbsp;</em></strong>The patent file is 1.7 million entries and can cause some delay in the command line, compared with smaller database files.</p> <p>Full details are available in this preprint by Barnabas et al (2022) DOI: <a href="https://doi.org/10.26434/chemrxiv-2022-nmnnd-v2">10.26434/chemrxiv-2022-nmnnd-v2</a></p> <p>Update 20/04/2022: uploaded files with updated CID mappings post-PubChem deposition.</p>

opencc-by-4.0Feb 2022View details →
zenodo48/100

Drug-patent DB

<p>A compressed file of a relational database of drug-related patented compounds (drug-patent DB). This DB was formatted using SQLite. The compound information was extracted from <a href="https://www.surechembl.org">SureChEMBL</a>.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Patents and certificates of addition granted in France from 1904 to 1921 by applicant's country

<p>**TITRE**<br> Etat des brevets d&#39;invention et des certificats d&#39;addition d&eacute;livr&eacute;s en France entre 1904 et 1921 par pays d&#39;origine du demandeur.</p> <p>**VARIABLES**</p> <p>YEAR : ann&eacute;e de d&eacute;livrance<br> COUNTRY : le nom donn&eacute; est le nom qui appara&icirc;t dans la source originale. Il n&#39;est simplifi&eacute; que pour quelques pays (Argentine, Luxembourg,...)<br> NUMBER : le nombre de brevets d&#39;invention et de certificats d&#39;addition d&eacute;livr&eacute;s.</p> <p>**OBSERVATIONS**<br> Pour l&#39;ann&eacute;e 1904, le nombre correspond au nombre de demandes de brevets et de certificats d&#39;addition</p> <p>**SOURCES**<br> ann&eacute;e 1904, <em>La Propri&eacute;t&eacute; industrielle</em>, avril 1906, p.&nbsp;59-60 ; ann&eacute;e 1905, <em>La Propri&eacute;t&eacute; industrielle</em>, mai 1907, p.&nbsp;74-76 ; ann&eacute;e 1906, <em>La Propri&eacute;t&eacute; industrielle</em>, juillet 1908, p 109-110 ; ann&eacute;e 1907, <em>Bulletin officiel de la propri&eacute;t&eacute; industrielle et commerciale</em>, 1908, p.&nbsp;32 ; ann&eacute;e 1908, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1910, p.&nbsp;13-14 ; ann&eacute;e 1909, <em>La Propri&eacute;t&eacute; industrielle</em>, mars 1911, p. 42-43 ; ann&eacute;e 1910, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1913, p.&nbsp;15-16 ; ann&eacute;e 1911, <em>La Propri&eacute;t&eacute; industrielle</em>, avril 1914, p.&nbsp;63-64 ; ann&eacute;e 1912, <em>Bulletin officiel de la propri&eacute;t&eacute; industrielle et commerciale</em>, 1913, p.&nbsp;23-24 ; ann&eacute;e 1913, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1915, p.&nbsp;11-12 ; ann&eacute;es 1914 et 1915, <em>La Propri&eacute;t&eacute; industrielle</em>, mai 1917, p.&nbsp;67-68 ; ann&eacute;e 1916, <em>La Propri&eacute;t&eacute; industrielle</em>, ao&ucirc;t 1918, p.&nbsp;95-96 ; ann&eacute;e 1917, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1919, p.&nbsp;11-12 ; ann&eacute;es 1918, 1919 et 1920, <em>La Propri&eacute;t&eacute; industrielle</em>, juin 1922, p.&nbsp;91-92 ; ann&eacute;e 1921, <em>Bulletin officiel de la propri&eacute;t&eacute; industrielle et commerciale</em>, 1922, p. 36.</p> <p>**TITLE**<br> Number of patents granted and certificates of addition in France between 1904 and 1921 by applicant&#39;s country</p> <p>**VARIABLES**</p> <p>YEAR : year of issue<br> COUNTRY : the name given is the one that appears in the original source. It is only simplified for a few countries (Argentina, Luxembourg,...).<br> NUMBER : the number of patents and certificates of addition granted.</p> <p>**OBSERVATIONS**<br> For the year 1904, the number corresponds to the number of applications for patents and certificates of addition.</p> <p>**SOURCES**<br> year 1904, <em>La Propri&eacute;t&eacute; industrielle</em>, avril 1906, p.&nbsp;59-60 ; year 1905, <em>La Propri&eacute;t&eacute; industrielle</em>, mai 1907, p.&nbsp;74-76 ; year 1906, <em>La Propri&eacute;t&eacute; industrielle</em>, juillet 1908, p 109-110 ; year 1907, <em>Bulletin officiel de la propri&eacute;t&eacute; industrielle</em>, 1908, p.&nbsp;32 ; year 1908, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1910, p.&nbsp;13-14 ; year 1909, <em>La Propri&eacute;t&eacute; industrielle</em>, mars 1911, p. 42-43 ; year 1910, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1913, p.&nbsp;15-16 ; year 1911, <em>La Propri&eacute;t&eacute; industrielle</em>, avril 1914, p.&nbsp;63-64 ; year 1912, <em>Bulletin officiel de la propri&eacute;t&eacute; industrielle</em> <em>et commerciale</em>, 1913, p.&nbsp;23-24 ; year 1913, <em>La Propri&eacute;t&eacute; industrielle,</em> janvier 1915, p.&nbsp;11-12 ; years 1914 et 1915, La Propri&eacute;t&eacute; industrielle, mai 1917, p.&nbsp;67-68 ; year 1916, <em>La Propri&eacute;t&eacute; industrielle</em>, ao&ucirc;t 1918, p.&nbsp;95-96 ; year 1917, <em>La Propri&eacute;t&eacute; industrielle</em>, janvier 1919, p.&nbsp;11-12 ; years 1918, 1919 et 1920, <em>La Propri&eacute;t&eacute; industrielle</em>, juin 1922, p.&nbsp;91-92 ; year 1921, <em>Bulletin officiel de la propri&eacute;t&eacute; industrielle et commerciale</em>, 1922, p. 36.&nbsp;&nbsp;&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Solicitantes de patente en Colombia (2000-2018) - Según tipo

<p>Relaci&oacute;n de las solicitudes de patente presentadas en colombia entre los a&ntilde;os 2000 y 2018, detallando cada uno de los solicitantes relacionados y si el mismo es nacional o extranjero</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Solicitantes de patente en Colombia (2000-2018) - Según tipo de persona

<p>Relaci&oacute;n de las solicitudes de patente presentadas en colombia entre los a&ntilde;os 2000 y 2018, detallando cada uno de los solicitantes relacionados y si el mismo es persona natural, empresa o universidad.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Países de prioridad asociados a las solicitudes de patente ante la WIPO que vinculen por lo menos a un colombiano

<p>Relaci&oacute;n de pa&iacute;ses de prioridad asociados con&nbsp;los registros de solicitudes de patente presentadas ante la OMPI entre los a&ntilde;os 2000 y 2019. que incluyen por lo menos a un colombiano como titular</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Sectores Tecnológicos asociados a las solicitudes de patente ante la WIPO que vinculen por lo menos a un colombiano

<p>Relaci&oacute;n de sectores tecnol&oacute;gicos asociados a los registros de solicitudes de patente presentadas ante la OMPI&nbsp;entre los a&ntilde;os 2000 y 2019.&nbsp;que incluyen por lo menos a un colombiano como titular</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

patccat: A classifier for patent claims

<p><strong>Data version: 3.3.0</strong></p> <p>Authors:<br>Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br>W. Keith Robinson (Wake Forest University, School of Law)<br>Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p><br>1. Notes on Data Construction<br>2. Citation and Code<br>3. Description of the Data Files<br>3.1. File List<br>3.2. List of Variables for Files with Claim-Level Information<br>3.3. List of Variables for Files with Patent-Level Information<br>4. Coming Soon!</p> <p><br><strong>1. Notes on Data Construction</strong></p> <p>This is version 3.3.0 of the patccat data (patent claim classification by algorithmic text analysis).</p> <p>Patent claims define an invention. A patent application is required to have one or more claims that distinctly claim the subject matter which the patent applicant regards as her invention or discovery. We construct a classifier of patent claims that identifies three distinct claim types: process claims, product claims, and product-by-process claims.</p> <p>For this classification, we combine information obtained from both the preamble and the body of a claim. The preamble is a general description of the invention (e.g., a method, an apparatus, or a device), whereas the body identifies steps and elements (specifying in detail the invention laid out in the preamble) that the applicant is claiming as the invention. The combination of the preamble type and the body type provides us with a more detailed and more accurate classification of claims than other approaches in the literature. This approach also accounts for unconventional drafting approaches. We eventually validate our classification using close to 10,000 manually classified claims.</p> <p>The data files contain the results of our classification. We provide claim-level information for each independent claim of U.S. utility patents granted between 1836 and 2020. We also provide patent-level information, i.e., the counts of different claim types for a given patent.</p> <p>For a detailed description of our classification approach, please take a look at the accompanying paper (Ganglmair, Robinson, and Seeligson 2022).</p> <p><strong>2. Citation</strong></p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): "The Rise of Process Claims: Evidence from a Century of U.S. Patents," unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p> <p>In the paper, we document the use of process claims in the U.S. over the last century, using the patccat data. We show an increase in the annual share of process claims of about 25 percentage points (from below 10% in 1920). This rise in process intensity of patents is not limited to a few patent classes, but we observe it across a broad spectrum of technologies. Process intensity varies by applicant type: companies file more process-intense patents than individuals, and U.S. applicants file more process-intense patents than foreign applicants. We further show that patents with higher process intensity are more valuable but are not necessarily cited more often. Last, process claims are on average shorter than product claims (with the gap narrowing since the 1970s).</p> <p>We would love to see how other researchers use the data and eventually learn from it. If you have a discussion paper or a publication in which you use the data, please send us a copy at patccat.data@gmail.com.</p> <p>We will the R code used to construct the data on Github with the next data version (version 3.4.0). Contact us at b.ganglmair@gmail.com if you would like to take a look at an earlier version of the code.</p> <p><br><strong>3. Description of the Data Files</strong></p> <p>The data files contain claim-level information for independent claims of 10,140,848 U.S. utility patents granted between 1836 and 2020. The files further contain patent-level information for U.S. utility patents.</p> <p><em>3.1. File List</em></p> File list <table><tbody> <tr> <td>claims-patccat-v3-3-sample.csv</td> <td>claim-level information for independent claims of a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>claims-patccat-v3-3-1836-1919.csv</td> <td>claim-level information for independent claims of 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>claims-patccat-v3-3-1920-2020.csv</td> <td>claim-level information for independent claims of 9,102,807 patents issued between 1920 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-sample.csv</td> <td>patent-level information for a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-1836-1919.csv</td> <td>patent-level information for 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>patents-patccat-v3-3-1920-2020.csv</td> <td>patent-level information for 9,102,807 patents issued between 1920 and 2020</td> </tr> </tbody> </table> <p><br><em>3.2. List of Variables for Files with Claim-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Claim-Level Information) <table><tbody> <tr> <td>PatentClaim</td> <td>patent claim identifier; 8-digit patent number and 4-digit claim number (Ex: 01234567-0001)</td> </tr> <tr> <td>singleLine</td> <td>=1 if claim is published in single-line format</td> </tr> <tr> <td>singleReformat</td> <td>outcome code of reformating of single-line claims</td> </tr> <tr> <td>Jepson</td> <td>=1 if claim is a Jepson claim</td> </tr> <tr> <td>JepsonReformat</td> <td>outcome code of reformating of Jepson claims</td> </tr> <tr> <td>inBegin</td> <td>=1 if claim begins with the word "in"</td> </tr> <tr> <td>wordsPreamble</td> <td>number of words in the claim preamble</td> </tr> <tr> <td>wordsBody</td> <td>number of words in the claim body</td> </tr> <tr> <td>dependentClaims</td> <td>number of dependent claims that refer to this independent claim</td> </tr> <tr> <td>isMeansPreamble</td> <td>=1 if term "means" is used in the preamble</td> </tr> <tr> <td>isMeansBody</td> <td>=1 if term "means" is used in the body</td> </tr> <tr> <td>isMeans</td> <td>=1 if term "means" is used anywhere in the claim (~ means-plus-function claim)</td> </tr> <tr> <td>processPreamble</td> <td>=1 if terms "method" or "process" are used in the preamble</td> </tr> <tr> <td>processBody</td> <td>=1 if terms "method" or "process" are used in the body</td> </tr> <tr> <td>processSimple</td> <td>=1 if terms "method" or "process" are used anywhere in the claim (for simple approach of process claim classification)</td> </tr> <tr> <td>claimType</td> <td>claim type of full classification (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>preambleType</td> <td>preamble type</td> </tr> <tr> <td>preambleTerm</td> <td>keyword used to classify preamble type</td> </tr> <tr> <td>preambleTermAlt</td> <td>alternative keyword (if preambleTerm were not used)</td> </tr> <tr> <td>preambleTextStub</td> <td>first 15 words of the preamble</td> </tr> <tr> <td>bodyType</td> <td>body type</td> </tr> <tr> <td>bodyLinesStep</td> <td>number of steps in the body</td> </tr> <tr> <td>bodyLinesElement</td> <td>number of elements in the body</td> </tr> <tr> <td>bodyLinesTotal</td> <td>total number of identified lines in the body</td> </tr> <tr> <td>label</td> <td>2-character label of the preamble-body combination; classification table maps label to claim type</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><em>3.3. List of Variables for Files with Patent-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Patent-Level Information) <table><tbody> <tr> <td>patent_id</td> <td>U.S. patent number (8-digit patent number)</td> </tr> <tr> <td>claims</td> <td>number of independent claims (the sum of the four claim types: 0, 1, 2, and 3)</td> </tr> <tr> <td>noCategory</td> <td>number of claims without a classified type</td> </tr> <tr> <td>processClaims</td> <td>number of process claims</td> </tr> <tr> <td>productClaims</td> <td>number of product claims</td> </tr> <tr> <td>prodByProcessClaims</td> <td>number of product-by-process claims</td> </tr> <tr> <td>firstClaim</td> <td>type of the first claim (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>simpleProcessClaims</td> <td>number of process claims by simple approach (terms "method" or "process" anywhere in the claim)</td> </tr> <tr> <td>simpleProcessPreamble</td> <td>number of process claims by simple approach (terms "method" or "process" in the preamble)</td> </tr> <tr> <td>meansClaims</td> <td>number of means-plus-function claims</td> </tr> <tr> <td>meansFirst</td> <td>=1 if first claim is a means-plus-function claim</td> </tr> <tr> <td>JepsonClaims</td> <td>number of Jepson claims</td> </tr> <tr> <td>JepsonFirst</td> <td>=1 if first claim is a Jepson claim</td> </tr> </tbody> </table> <p><br>Note: The following variables/fields are currently empty (March 30, 2020); we will populate these variables/fields with data version 3.4.0.</p> <p>preambleTerm<br>preambleTermAlt<br>preambleTextStub<br>bodyLinesStep<br>bodyLinesElement<br>bodyLinesTotal</p> <p>Note: We will release the data for patents issued in 2021 with data version 3.4.0.</p> <p><br><strong>4. Coming Soon!</strong></p> <p>We are working on a number of extensions of the patccat data.</p> <p>- With data version 3.4.0, we plan to release data for all published U.S. patent applications (2001 through 2021)<br>- In late spring/early summer 2022, we will release data for patents issued by the European Patent Office (EPO) [<strong>Update: March 28, 2023</strong>: see <a href="https://doi.org/10.5281/zenodo.7776092">https://doi.org/10.5281/zenodo.7776092</a>]<br>- In late spring/early summer 2022, we will release data for patents issued by the Canadian Intellectual Property Office (CIPO)</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

TecKnoGraph: Knowledge Graph from patents in C4ISTAR

<p>&lt;img src=&quot;https://github.com/nicolamelluso/TecKnoGraph-demo/blob/main/TecKnoGraph-Example%20Graph.png&quot; alt=&quot;TecKnoGraph&quot;&gt;</p> <p>This dataset contains a sample of Knowledge Graph (KG) created with TecKnoGraph.</p> <p>There are two files:</p> <p><strong>- TecKnoGraph-C4ISTAR-sample.csv</strong>: this file&nbsp;contains the KG in the form of triples where each element of the triple (source, relation, target) is tagged with categories.</p> <p><strong>- patents.zip:</strong>&nbsp;this file contains data about 10,000 patents; each patent corresponds to a txt file.</p> <p>There is available a demo for using TecKnoGraph from examples of input text:<br> https://nicolamelluso-tecknograph-demo-tecknograph-streamlit-cnq4sm.streamlitapp.com/</p> <p>In this repository it is possible to find also the appendix of the corresponding paper.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis

<p>Dataset of the research paper:&nbsp;<strong>Impact of Software Engineering Research in Practice:&nbsp;A Patent and Author Survey Analysis</strong></p> <p>Existing work on the practical impact of software engineering (SE) research examines industrial relevance rather than adoption of study results, hence the question of how results have been practically applied remains open. To answer this and investigate the outcomes of impactful research, we performed a quantitative and qualitative analysis of 4 354 SE patents citing 1 690 SE papers published in four leading SE venues between 1975&ndash;2017. Moreover, we conducted a survey on 475 authors of 593 top-cited and awarded publications, achieving 26% response rate. Overall, researchers have equipped practitioners with various tools, processes, and methods, and improved many existing products. SE practice values knowledge-seeking research and is impacted by diverse cross-disciplinary SE areas. Practitioner-oriented publication venues appear more impactful than researcher-oriented ones, while industry-related tracks in conferences could enhance their impact. Some research works did not reach a wide footprint due to limited funding resources or unfavorable cost-benefit trade-off of the proposed solutions. The need for higher SE research funding could be corroborated through a dedicated empirical study. In general, the assessment of impact is subject to its definition. Therefore, academia and industry could jointly agree on a formal description to set a common ground for subsequent research on the topic.</p> <p>The following data&nbsp;files are included.</p> <ul> <li><em>./fields</em>: <ul> <li><strong>engi-fields.csv</strong>: Publication and PhD dissertation counts of main engineering branches</li> <li><strong>engi-fields-queries.txt</strong>: Queries applied to Elsevier&#39;s Scopus and Open Access Theses and Dissertations databases to retrieve the publication and dissertation counts</li> </ul> </li> <li><em>./patents</em>: <ul> <li><strong>sample-se-references-verified.csv</strong>: Manual verification of a random sample of references by software engineering (SE) patents to SE papers</li> <li><strong>se-cpc.tsv</strong>: Manually-identified SE-related Cooperative Patent Classification (CPC) categories</li> <li><strong>se-references-in-patents.csv</strong>: SE references made by SE patents to SE papers</li> <li><em>./patents/litigation</em>: <ul> <li><strong>case-values.csv</strong>: Manually-retrieved litigation damages of citing SE patents</li> <li><strong>lit-per-paper.csv</strong>: Litigation cases of citing SE patents</li> </ul> </li> <li><em>./patents/maintenance</em>: <ul> <li><strong>maint-code-fee-mapping.csv</strong>: Mapping of patent maintenance fee codes to their fee values</li> <li><strong>maint-fees.csv</strong>: Fee values of maintenance fee codes</li> <li><strong>maint-per-paper.csv</strong>: Maintenance fee events of citing SE patents</li> </ul> </li> <li><em>./patents/reports</em>: <ul> <li><strong>lit-sum-per-paper.csv</strong>: Counts and total damages of litigation cases of patent-cited SE papers</li> <li><strong>maint-sum-per-paper.csv</strong>: Counts and total values of maintenance fee events of patent-cited SE papers</li> <li><strong>patent-ref-counts.csv</strong>: SE patent citation counts of patent-cited SE papers</li> </ul> </li> </ul> </li> <li><em>./survey</em>: <ul> <li><strong>emse-top.csv</strong>: Most-cited papers of the Empirical Software Engineering (EMSE) journal</li> <li><strong>icse-bp.csv</strong>: Distinguished papers of the International Conference of Software Engineering (ICSE)</li> <li><strong>icse-mip.csv</strong>: Most influential ICSE papers</li> <li><strong>icse-top.csv</strong>: Most-cited ICSE papers</li> <li><strong>survey-questionnaire-emse.pdf</strong>: The EMSE survey questionnaire</li> <li><strong>survey-questionnaire.pdf</strong>: The ICSE, TSE, and TOSEM&nbsp;survey questionnaire</li> <li><strong>survey-responses.csv</strong>: The anonymized survey responses</li> <li><strong>tosem-top.csv</strong>: Most-cited papers of the ACM Transactions on Software Engineering and Methodology (TOSEM)</li> <li><strong>tse-top.csv</strong>: Most-cited papers of the IEEE Transactions on Software Engineering (TSE)</li> <li><em>./survey/manual-coding</em>: <ul> <li><strong>feedback.txt</strong>: Manual coding of survey feedback</li> <li><strong>practical-impact.csv</strong>: Manual coding of responses about practical impact of work</li> <li><strong>practical-impact-lack.csv</strong>: Manual coding of responses about lack of practical impact</li> <li><strong>research-methods.csv</strong>: Manual coding of additional research methods of surveyed papers</li> <li><strong>state-of-practice.csv</strong>: Manual coding of responses about changes in state of practice</li> </ul> </li> </ul> </li> <li><em>./venues</em>: <ul> <li><strong>se-venues.csv</strong>: Top SE venues according to Google Scholar Metrics</li> <li><strong>se-venues-impact.csv</strong>: SE patent citations and patent-based impact factors of SE venues</li> <li><strong>se-venues-scopus-queries.txt</strong>: Queries applied to Scopus to retrieve the publication counts of the SE venues</li> </ul> </li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo44/100

A Patent Survey on the Biotechnological Production of 2,5-Furandicarboxylic acid (FDCA): Current Trends and Challenges

<p>The production of 2,5-furandicarboxylic acid (FDCA) as a biobased commodity chemical has gained great importance over the last years. The possibility of replacing conventional polyethylene terephthalate (PET)-based plastics by polyethylene furanoate (PEF) is accelerating applied research in the direction of novel and highly productive synthesis routes to FDCA. This paper explores the patent activities related to FDCA production, with particular emphasis on the potential role that enzymatic catalytic methods may have. An increasing number of patent applications (and granted ones) have been disclosed over the last decade. The innovation is set on the development of multi-step catalytic processes, involving multi-enzymatic and chemoenzymatic cascades, or using whole-cells, either as resting (non-growing) biocatalysts or as living organisms in fermentative processes. Moreover, other innovative paths use substrates different from fructose and 5-hydroxymethylfurfural (HMF), such as gluconic acid, or furfural. Overall, opportunities for innovation exist, albeit current production metrics in biotechnological methods remain mostly at the proof-of-concept level. Large development efforts to reach industrial targets are needed, which may be stimulated by the less severe processing conditions expected for biotechnology, leading to energy savings, less by-product formation, and to the valorization of crude effluents.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

MADIA_732678_INN_patents_02

<p>Collection of .csv files with raw patent data, analysed to obtain landscaping information on technologies overlapping with the project. Used to produce the deliverable &ldquo;Patent and scientific literature study M24&rdquo; D7.6 Patent landscaping. The data where obtained on <a href="http://www.thelens.org">www.thelens.org</a></p> <p>The queries used to obtain the data were:</p> <p>A</p> <p>title:(neurodegenerative early diagnosis) OR abstract:(neurodegenerative early diagnosis) OR claims:(neurodegenerative early diagnosis) AND classification_ipcr:((G01N*) OR (H01L*))</p> <p>B</p> <p>title:(functionalized magnetic nano*) OR abstract:(functionalized magnetic nano*) OR claims: (functionalized magnetic nano*) AND classification_ipcr:((G01N*) OR (B82Y*))</p> <p>C</p> <p>title:((microfluidic) AND (chamber OR apparatus OR device)) OR abstract:((microfluidic) AND (chamber OR apparatus OR device)) OR claims:((microfluidic) AND (chamber OR apparatus OR device)) AND classification_ipcr:((B01J*) OR (B81B*))</p> <p>D</p> <p>title:((magnetic nano*) AND (sensor OR detector)) OR abstract:((magnetic nano*) AND (sensor OR detector)) OR claims: ((magnetic nano*) AND (sensor OR detector)) AND classification_ipcr:((G01R*) OR (G11B*) OR (H01L*) OR (H01F*))</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

3PFL: Database of Patents and Publications with a Public-Funding Linkage

<p>The 3PFL database links information on patented inventions and scientific publications related to a public procurement contract or a research grant awarded by the U.S. Federal Government to detailed contract-level/grant-level information (e.g., awarding agency, recipient organization, award size). We have combined data from multiple sources, including (but not limited to) the United States Patent and Trademark Office bulk database, the Federal Procurement Database System, the Award Submission Portal (ASP), and the European Patent Office&#39;s PATSTAT database.&nbsp;We also provide a link to the scientific publications associated with these patents. The 3PFL database provides rich and original information that opens the door to novel empirical research in the economics of innovation and science.&nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

AI-related patents (WIPO, category G06N) and market capitalisation by companies registering at least 2 new ones in 2019, sorted into four global regions (China, USA, EEA, rest of the world)

<p>NOTE: for some reason the pptx and previews keep getting munged on this supposedly permanent arxiv, but the data is still there, unchanged, and you can see how the pptx should look in either the jpg, or the the article.</p> <p>Datasets and presentations concerning the strength of the EU and "the rest of the world" relative to China and&nbsp;the USA, for the purpose of illustrating and counteracting / better informing narratives concerning a "new AI cold war". The materials authored by us may be freely used under the terms of the MIT License, which appears in its entirety in both the dataset and the presentation. The other materials are only curated by us, taken from Twitter as examples of misinformation pertaining to this concern.</p> <p>As of 28 June, this work now also appears in a formal publication:&nbsp;Joanna J. Bryson, Helena Malikova; Is There an AI Cold War?.&nbsp;<em><em>Global Perspectives</em></em>&nbsp;2021; 2 (1): 24803. doi:&nbsp;<a href="https://doi.org/10.1525/gp.2021.24803">https://doi.org/10.1525/gp.2021.24803</a></p> <p>Authors: The original analysis was conducted primarily by Malikova in collaboration with Bryson. An associated publication is anticipated where Bryson is the lead author.</p> <p>Contributors: independently followed Malikova's procedures to check her work.&nbsp;Inconsistencies were triple checked and resolved.</p>

openmit-licenseOct 2020View details →
zenodo44/100

The patccat classifier for patent claims - EPO edition

<p>!!! This is the <strong>EPO/European version</strong> of the patccat classifier of patent claims. !!!</p> <p>Note: We use the same approach that we use for USPTO patents. For a detailed description, see <a href="https://doi.org/10.5281/zenodo.6395307">https://doi.org/10.5281/zenodo.6395307</a>.</p> <p><strong>Data version: 3.4.0</strong></p> <p>Authors:<br> Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br> W. Keith Robinson (Wake Forest University, School of Law)<br> Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): &quot;The Rise of Process Claims: Evidence from a Century of U.S. Patents,&quot; unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Patent text: code, data, and new measures

<p>This Zenodo page describes data collection, processing, and different open access data files related to the text of USPTO patent documents. The document &quot;Data Description Zenodo.pdf&quot;&nbsp;provides more details.&nbsp;If you use the code or data, please cite the following paper:</p> <p>Arts&nbsp;S, Hou&nbsp;J,&nbsp;Gomez&nbsp;JC&nbsp;(2021). Natural language processing to identify the creation and impact of new technologies in patent text: Code, data, and new measures. <em>Research Policy</em>, 50(2), 104144.&nbsp;(<a href="https://doi.org/10.1016/j.respol.2020.104144">https://doi.org/10.1016/j.respol.2020.104144</a>)</p>

opencc-by-nc-1.0Nov 2020View details →
zenodo40/100

PatCit: A Comprehensive Dataset of Patent Citations

<p><em><strong>patCit:&nbsp;A Comprehensive Dataset of Patent Citations</strong></em>&nbsp;[<a href="https://tinyletter.com/patcit">Newsletter</a>,&nbsp;<a href="https://github.com/cverluise/PatCit">GitHub</a>]</p> <p>Patents are at the crossroads of many innovation nodes: science, industry, products, competition, etc. Such interactions can be identified through citations&nbsp;<em>in a broad sense</em>.</p> <p>It is now common to use front-page patent citations to study some aspects of the innovation system. However, <strong>there is much more buried in the Non Patent Literature (NPL) citations and in the patent text itself</strong>.&nbsp;<strong>patCit extracts and structures these citations.</strong></p> <blockquote> <p>Want to know more? Read patCit&nbsp;<a href="https://docs.google.com/presentation/d/11COlz64EZn8PipXvnDBBZI_bnDD0fpm6tyx1_EqD6lU/edit?usp=sharing">academic presentation</a>&nbsp;or dive into usage and technical guides on patCit&nbsp;<a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p> </blockquote> <p><strong><em>IN PRACTICE</em></strong></p> <p>At patCit, we are building a&nbsp;<em>comprehensive</em>&nbsp;dataset of patent citations to help the community explore this&nbsp;<em>terra incognita</em>. patCit has the following features:</p> <ul> <li>global coverage</li> <li>front-page and in-text citations</li> <li>all categories&nbsp;of NPL documents</li> </ul> <p><strong><em>Front-page</em></strong></p> <p>patCit builds on&nbsp;<a href="https://www.epo.org/searching-for-patents/data/bulk-data-sets/docdb.html#tab-1">DOCDB</a>, the largest database of Non Patent Literature (NPL) citations. First, we deduplicate this corpus and organize it into 10 categories (bibliographical reference, database, norm &amp; standard, etc). Then, we design and apply category specific information extraction models using&nbsp;<a href="https://github.com/explosion/spaCy">spaCy</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p><strong><em>In-text</em></strong></p> <p>patCit builds on Google Patents corpus of&nbsp;<a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&amp;p=patents-public-data&amp;d=patents&amp;t=publications&amp;page=table">USPTO full-text patents</a>. First, we extract patent and bibliographical reference citations. Then, we parse detected in-text citations into a series of category dependent attributes using&nbsp;<a href="https://github.com/kermitt2/grobid">grobid</a>. Patent citations are matched with a standard publication number using the Google Patents&nbsp;<a href="https://patents.google.com/api/match">matching API</a>&nbsp;and bibliographical references are matched with a DOI using&nbsp;<a href="https://github.com/kermitt2/biblio-glutton">biblio-glutton</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p>&nbsp;</p> <p><strong>FAIR</strong></p> <p><strong>Find</strong>&nbsp;- The patCit dataset is available on&nbsp;<a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&amp;p=patcit-public-data&amp;page=project">BigQuery</a>&nbsp;in an interactive environment. For those who have a smattering of SQL, this is the perfect place to explore the data. It can also be downloaded on&nbsp;<a href="https://zenodo.org/record/3710994">Zenodo</a>.</p> <p><strong>Interoperate</strong>&nbsp;- Interoperability is at the core of patCit ambition. We take care to extract unique identifiers whenever it is possible to enable data enrichment for domain specific high quality databases. This includes the DOI, PMID and PMCID for bibliographical references, the Technical Doc Number for standards, the Accession Number for Genetic databases, the publication number for PATSTAT and Claims, etc. See specific table for more details.</p> <p><strong>Reproduce</strong>&nbsp;- Our <a href="https://github.com/cverluise/PatCit">gitHub</a> repository is the project factory. You can learn more about data recipes and models on the patCit&nbsp;<a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

USPTO patent data: 250k random sample and NPEs' patents

<p>This package includes:</p> <ol> <li>Two Stata .dta files consisting of information on patents assigned by the United States Patent and Trademark Office between 1976 and 2014: a random sample of 250 000 US patents, and data on patent owned by Intellectual Ventures, RPX, and several other companies. The variables for example include: grant date, application date, forward and backward citations, renewals, claims&nbsp;and others.</li> <li>Source codes and methods used in generating and analyzing the two data files.</li> </ol> <p>A bachelor thesis with further information will be linked here.</p>

opencc-zeroMay 2016View details →
zenodo40/100

Link Compustat – USPTO Patent Assignment Dataset

<p>This page provides the data resulting from linking assignees and assignors in the USPTO Patent Assignment Dataset to Compustat gvkeys. We work with a version of the USPTO PAD that was gracefully shared with us by Stuart Graham. Such version precedes by one year the first release available at the USPTO website (https://www.uspto.gov/ip-policy/economic-research/research-datasets/patent-assignment-dataset). The version that we use covers 5,534,135 transactions recorded at the USPTO between January 1970 and January 2013 (inclusive). While the first transaction date is January 1970, the number of transactions recorded in the initial years is negligible. Data coverage seems sufficient for the years 1981-2012.</p> <p>&nbsp;</p> <p>If you use the code or data, please cite the following two papers:</p> <p>&nbsp;</p> <p>Arque-Castells, P., and Spulber, D. (2022). Measuring the Private and Social Returns to R&amp;D: Unintended Spillovers versus Technology Markets. Journal of Political Economy. <a href="https://doi.org/10.1086/719908">https://doi.org/10.1086/719908</a></p> <p>&nbsp;</p> <p>Arqu&eacute; Castells, Pere and Spulber, Daniel F., Firm Matching in the Market for Technology: Business Stealing and Business Creation (September 17, 2021). Northwestern Law &amp; Econ Research Paper No. 18-14, Available at SSRN: https://ssrn.com/abstract=3041558 or <a href="http://dx.doi.org/10.2139/ssrn.3041558">http://dx.doi.org/10.2139/ssrn.3041558</a></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record