Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
267
datasets available to search
ShareScore release 0.9.0
Dataset results
267 results for “Patent”
Data for: 3D bioprinting patents
<p>This dataset contains information regarding 3D bioprinting patent/patent applications. </p> <p><a href="https://www.orbit.com/">Orbit</a> (a fee-based patent database provided by Questel), accessed on Jan. 2, 2022, was used for data mining.</p> <p>The files titled “<em>3D bioprinting patents</em>” and “<em>Bioink patents</em>” contain information related to Priority, Application and Publication numbers, Priority Application and Publication dates, Title, Abstract, and Current assignees.</p> <p>The patent searches were carried out by keywords and classification codes. </p> <p>Both IPC (International Patent Classification) and CPC (Cooperative Patent Classification) codes were used. </p> <p>Instead of 309 patents (see reference 1), a total number of 3,681 documents were retrieved (of which 3,027 are still alive and 2,461 filed in the period 2016 – 2020).</p> <p>Most of the published patent applications were generated in China and the USA (1,360 vs 1,063 priority applications). China is effectively the leading country (37%), followed by the USA with 29% of priority patent applications. </p> <p><strong>Value of the dataset</strong>: prior art searches, technological trends </p> <p><strong>Steps to reproduce data</strong>: </p> <p>The search strategy is reported in the table below:</p> <table> <tbody> <tr> <td> <p>1</p> </td> <td> <p>857</p> </td> <td> <p>(BIOPRINT+ OR BIOINK? OR ORGAN_ON_A_CHIP)/TI/AB/CLMS/ICLM</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>3096</p> </td> <td> <p>((A61L-027+ OR A61F-002+ OR A61L2430/00)</p> <p>AND (B33Y+ OR B29C-064+))/IPC/CPC</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>3681</p> </td> <td> <p> 1 OR 2</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>71912</p> </td> <td> <p>(C09D-011+)/IPC/CPC</p> </td> </tr> <tr> <td> <p>5</p> </td> <td> <p>104</p> </td> <td> <p> 3 AND 4</p> </td> </tr> </tbody> </table> <p><strong>Definition of the classification codes used</strong></p> <p><strong>A61L 27</strong>: Materials for grafts or prostheses or for coating grafts or prostheses</p> <p><strong>A61F 2</strong>: Filters implantable into blood vessels; Prostheses, i.e., artificial substitutes or replacements for parts of the body; Appliances for connecting them with the body; Devices providing patency to, or preventing collapsing of, tubular structures of the body, e.g., stents</p> <p><strong>A61L2430/00</strong>: Materials or treatment for tissue regeneration</p> <p><strong>B33Y</strong>: Additive manufacturing, i.e., manufacturing of three-dimensional [3-d] objects by additive deposition, additive agglomeration, or additive layering, e.g., by 3-d printing, stereolithography, or selective laser sintering</p> <p><strong>B29C 64/00</strong>: Additive manufacturing, i.e., manufacturing of three-dimensional [3D] objects by additive deposition, additive agglomeration, or additive layering, e.g., by 3D printing, stereolithography, or selective laser sintering</p> <p><strong>C09D 11</strong>: inks </p>
OntoChem PFAS CORE and Patent Files for MetFrag
<p>These are MetFraggable versions of the PFAS lists produced by OntoChem by performing literature mining of the CORE database (27K entries) and Google Patent collections (1.7M entries). The MetFrag versions have undergone filtering to remove entries that prevent MetFrag running (i.e. multiple-entry formulas and certain elements). Each file contains a tag whether the PFAS fits three definitions, A, B and C. The CORE database also contains the number of references in which the PFAS entry was found. Each file is available as CSV or in compressed form.</p> <ul> <li>PFAS Definition A: Each compound that contains a CF<sub>2</sub> group</li> <li>PFAS Definition B: Each compound that contains a (AH)(AH)(F)C-C(AH)F<sub>2</sub> group, where AH groups could be hydrogen or any other atom and the bond between both aliphatic carbon atoms is a single bond</li> <li>PFAS Definition C: Each compound that contains a (R<sup>1</sup>)(R<sup>2</sup>)(F)C-C(R<sup>3</sup>)F<sub>2</sub> group is considered a PFAS, where the R groups are any atom except hydrogen and the bond between both aliphatic carbon atoms is a single bond</li> </ul> <p><strong><em>Please note these files are very large (especially patents) and should not be uploaded to MetFragWeb directly - they will be available from the dropdown menu. These are provided for command line users, and any other workflows interested in these files! </em></strong>The patent file is 1.7 million entries and can cause some delay in the command line, compared with smaller database files.</p> <p>Full details are available in this preprint by Barnabas et al (2022) DOI: <a href="https://doi.org/10.26434/chemrxiv-2022-nmnnd-v2">10.26434/chemrxiv-2022-nmnnd-v2</a></p> <p>Update 20/04/2022: uploaded files with updated CID mappings post-PubChem deposition.</p>
Drug-patent DB
<p>A compressed file of a relational database of drug-related patented compounds (drug-patent DB). This DB was formatted using SQLite. The compound information was extracted from <a href="https://www.surechembl.org">SureChEMBL</a>.</p>
Patents and certificates of addition granted in France from 1904 to 1921 by applicant's country
<p>**TITRE**<br> Etat des brevets d'invention et des certificats d'addition délivrés en France entre 1904 et 1921 par pays d'origine du demandeur.</p> <p>**VARIABLES**</p> <p>YEAR : année de délivrance<br> COUNTRY : le nom donné est le nom qui apparaît dans la source originale. Il n'est simplifié que pour quelques pays (Argentine, Luxembourg,...)<br> NUMBER : le nombre de brevets d'invention et de certificats d'addition délivrés.</p> <p>**OBSERVATIONS**<br> Pour l'année 1904, le nombre correspond au nombre de demandes de brevets et de certificats d'addition</p> <p>**SOURCES**<br> année 1904, <em>La Propriété industrielle</em>, avril 1906, p. 59-60 ; année 1905, <em>La Propriété industrielle</em>, mai 1907, p. 74-76 ; année 1906, <em>La Propriété industrielle</em>, juillet 1908, p 109-110 ; année 1907, <em>Bulletin officiel de la propriété industrielle et commerciale</em>, 1908, p. 32 ; année 1908, <em>La Propriété industrielle</em>, janvier 1910, p. 13-14 ; année 1909, <em>La Propriété industrielle</em>, mars 1911, p. 42-43 ; année 1910, <em>La Propriété industrielle</em>, janvier 1913, p. 15-16 ; année 1911, <em>La Propriété industrielle</em>, avril 1914, p. 63-64 ; année 1912, <em>Bulletin officiel de la propriété industrielle et commerciale</em>, 1913, p. 23-24 ; année 1913, <em>La Propriété industrielle</em>, janvier 1915, p. 11-12 ; années 1914 et 1915, <em>La Propriété industrielle</em>, mai 1917, p. 67-68 ; année 1916, <em>La Propriété industrielle</em>, août 1918, p. 95-96 ; année 1917, <em>La Propriété industrielle</em>, janvier 1919, p. 11-12 ; années 1918, 1919 et 1920, <em>La Propriété industrielle</em>, juin 1922, p. 91-92 ; année 1921, <em>Bulletin officiel de la propriété industrielle et commerciale</em>, 1922, p. 36.</p> <p>**TITLE**<br> Number of patents granted and certificates of addition in France between 1904 and 1921 by applicant's country</p> <p>**VARIABLES**</p> <p>YEAR : year of issue<br> COUNTRY : the name given is the one that appears in the original source. It is only simplified for a few countries (Argentina, Luxembourg,...).<br> NUMBER : the number of patents and certificates of addition granted.</p> <p>**OBSERVATIONS**<br> For the year 1904, the number corresponds to the number of applications for patents and certificates of addition.</p> <p>**SOURCES**<br> year 1904, <em>La Propriété industrielle</em>, avril 1906, p. 59-60 ; year 1905, <em>La Propriété industrielle</em>, mai 1907, p. 74-76 ; year 1906, <em>La Propriété industrielle</em>, juillet 1908, p 109-110 ; year 1907, <em>Bulletin officiel de la propriété industrielle</em>, 1908, p. 32 ; year 1908, <em>La Propriété industrielle</em>, janvier 1910, p. 13-14 ; year 1909, <em>La Propriété industrielle</em>, mars 1911, p. 42-43 ; year 1910, <em>La Propriété industrielle</em>, janvier 1913, p. 15-16 ; year 1911, <em>La Propriété industrielle</em>, avril 1914, p. 63-64 ; year 1912, <em>Bulletin officiel de la propriété industrielle</em> <em>et commerciale</em>, 1913, p. 23-24 ; year 1913, <em>La Propriété industrielle,</em> janvier 1915, p. 11-12 ; years 1914 et 1915, La Propriété industrielle, mai 1917, p. 67-68 ; year 1916, <em>La Propriété industrielle</em>, août 1918, p. 95-96 ; year 1917, <em>La Propriété industrielle</em>, janvier 1919, p. 11-12 ; years 1918, 1919 et 1920, <em>La Propriété industrielle</em>, juin 1922, p. 91-92 ; year 1921, <em>Bulletin officiel de la propriété industrielle et commerciale</em>, 1922, p. 36. </p>
Solicitantes de patente en Colombia (2000-2018) - Según tipo
<p>Relación de las solicitudes de patente presentadas en colombia entre los años 2000 y 2018, detallando cada uno de los solicitantes relacionados y si el mismo es nacional o extranjero</p>
Solicitantes de patente en Colombia (2000-2018) - Según tipo de persona
<p>Relación de las solicitudes de patente presentadas en colombia entre los años 2000 y 2018, detallando cada uno de los solicitantes relacionados y si el mismo es persona natural, empresa o universidad.</p>
Países de prioridad asociados a las solicitudes de patente ante la WIPO que vinculen por lo menos a un colombiano
<p>Relación de países de prioridad asociados con los registros de solicitudes de patente presentadas ante la OMPI entre los años 2000 y 2019. que incluyen por lo menos a un colombiano como titular</p>
Sectores Tecnológicos asociados a las solicitudes de patente ante la WIPO que vinculen por lo menos a un colombiano
<p>Relación de sectores tecnológicos asociados a los registros de solicitudes de patente presentadas ante la OMPI entre los años 2000 y 2019. que incluyen por lo menos a un colombiano como titular</p>
patccat: A classifier for patent claims
<p><strong>Data version: 3.3.0</strong></p> <p>Authors:<br>Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br>W. Keith Robinson (Wake Forest University, School of Law)<br>Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p><br>1. Notes on Data Construction<br>2. Citation and Code<br>3. Description of the Data Files<br>3.1. File List<br>3.2. List of Variables for Files with Claim-Level Information<br>3.3. List of Variables for Files with Patent-Level Information<br>4. Coming Soon!</p> <p><br><strong>1. Notes on Data Construction</strong></p> <p>This is version 3.3.0 of the patccat data (patent claim classification by algorithmic text analysis).</p> <p>Patent claims define an invention. A patent application is required to have one or more claims that distinctly claim the subject matter which the patent applicant regards as her invention or discovery. We construct a classifier of patent claims that identifies three distinct claim types: process claims, product claims, and product-by-process claims.</p> <p>For this classification, we combine information obtained from both the preamble and the body of a claim. The preamble is a general description of the invention (e.g., a method, an apparatus, or a device), whereas the body identifies steps and elements (specifying in detail the invention laid out in the preamble) that the applicant is claiming as the invention. The combination of the preamble type and the body type provides us with a more detailed and more accurate classification of claims than other approaches in the literature. This approach also accounts for unconventional drafting approaches. We eventually validate our classification using close to 10,000 manually classified claims.</p> <p>The data files contain the results of our classification. We provide claim-level information for each independent claim of U.S. utility patents granted between 1836 and 2020. We also provide patent-level information, i.e., the counts of different claim types for a given patent.</p> <p>For a detailed description of our classification approach, please take a look at the accompanying paper (Ganglmair, Robinson, and Seeligson 2022).</p> <p><strong>2. Citation</strong></p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): "The Rise of Process Claims: Evidence from a Century of U.S. Patents," unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p> <p>In the paper, we document the use of process claims in the U.S. over the last century, using the patccat data. We show an increase in the annual share of process claims of about 25 percentage points (from below 10% in 1920). This rise in process intensity of patents is not limited to a few patent classes, but we observe it across a broad spectrum of technologies. Process intensity varies by applicant type: companies file more process-intense patents than individuals, and U.S. applicants file more process-intense patents than foreign applicants. We further show that patents with higher process intensity are more valuable but are not necessarily cited more often. Last, process claims are on average shorter than product claims (with the gap narrowing since the 1970s).</p> <p>We would love to see how other researchers use the data and eventually learn from it. If you have a discussion paper or a publication in which you use the data, please send us a copy at patccat.data@gmail.com.</p> <p>We will the R code used to construct the data on Github with the next data version (version 3.4.0). Contact us at b.ganglmair@gmail.com if you would like to take a look at an earlier version of the code.</p> <p><br><strong>3. Description of the Data Files</strong></p> <p>The data files contain claim-level information for independent claims of 10,140,848 U.S. utility patents granted between 1836 and 2020. The files further contain patent-level information for U.S. utility patents.</p> <p><em>3.1. File List</em></p> File list <table><tbody> <tr> <td>claims-patccat-v3-3-sample.csv</td> <td>claim-level information for independent claims of a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>claims-patccat-v3-3-1836-1919.csv</td> <td>claim-level information for independent claims of 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>claims-patccat-v3-3-1920-2020.csv</td> <td>claim-level information for independent claims of 9,102,807 patents issued between 1920 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-sample.csv</td> <td>patent-level information for a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-1836-1919.csv</td> <td>patent-level information for 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>patents-patccat-v3-3-1920-2020.csv</td> <td>patent-level information for 9,102,807 patents issued between 1920 and 2020</td> </tr> </tbody> </table> <p><br><em>3.2. List of Variables for Files with Claim-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Claim-Level Information) <table><tbody> <tr> <td>PatentClaim</td> <td>patent claim identifier; 8-digit patent number and 4-digit claim number (Ex: 01234567-0001)</td> </tr> <tr> <td>singleLine</td> <td>=1 if claim is published in single-line format</td> </tr> <tr> <td>singleReformat</td> <td>outcome code of reformating of single-line claims</td> </tr> <tr> <td>Jepson</td> <td>=1 if claim is a Jepson claim</td> </tr> <tr> <td>JepsonReformat</td> <td>outcome code of reformating of Jepson claims</td> </tr> <tr> <td>inBegin</td> <td>=1 if claim begins with the word "in"</td> </tr> <tr> <td>wordsPreamble</td> <td>number of words in the claim preamble</td> </tr> <tr> <td>wordsBody</td> <td>number of words in the claim body</td> </tr> <tr> <td>dependentClaims</td> <td>number of dependent claims that refer to this independent claim</td> </tr> <tr> <td>isMeansPreamble</td> <td>=1 if term "means" is used in the preamble</td> </tr> <tr> <td>isMeansBody</td> <td>=1 if term "means" is used in the body</td> </tr> <tr> <td>isMeans</td> <td>=1 if term "means" is used anywhere in the claim (~ means-plus-function claim)</td> </tr> <tr> <td>processPreamble</td> <td>=1 if terms "method" or "process" are used in the preamble</td> </tr> <tr> <td>processBody</td> <td>=1 if terms "method" or "process" are used in the body</td> </tr> <tr> <td>processSimple</td> <td>=1 if terms "method" or "process" are used anywhere in the claim (for simple approach of process claim classification)</td> </tr> <tr> <td>claimType</td> <td>claim type of full classification (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>preambleType</td> <td>preamble type</td> </tr> <tr> <td>preambleTerm</td> <td>keyword used to classify preamble type</td> </tr> <tr> <td>preambleTermAlt</td> <td>alternative keyword (if preambleTerm were not used)</td> </tr> <tr> <td>preambleTextStub</td> <td>first 15 words of the preamble</td> </tr> <tr> <td>bodyType</td> <td>body type</td> </tr> <tr> <td>bodyLinesStep</td> <td>number of steps in the body</td> </tr> <tr> <td>bodyLinesElement</td> <td>number of elements in the body</td> </tr> <tr> <td>bodyLinesTotal</td> <td>total number of identified lines in the body</td> </tr> <tr> <td>label</td> <td>2-character label of the preamble-body combination; classification table maps label to claim type</td> </tr> </tbody> </table> <p> </p> <p><em>3.3. List of Variables for Files with Patent-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Patent-Level Information) <table><tbody> <tr> <td>patent_id</td> <td>U.S. patent number (8-digit patent number)</td> </tr> <tr> <td>claims</td> <td>number of independent claims (the sum of the four claim types: 0, 1, 2, and 3)</td> </tr> <tr> <td>noCategory</td> <td>number of claims without a classified type</td> </tr> <tr> <td>processClaims</td> <td>number of process claims</td> </tr> <tr> <td>productClaims</td> <td>number of product claims</td> </tr> <tr> <td>prodByProcessClaims</td> <td>number of product-by-process claims</td> </tr> <tr> <td>firstClaim</td> <td>type of the first claim (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>simpleProcessClaims</td> <td>number of process claims by simple approach (terms "method" or "process" anywhere in the claim)</td> </tr> <tr> <td>simpleProcessPreamble</td> <td>number of process claims by simple approach (terms "method" or "process" in the preamble)</td> </tr> <tr> <td>meansClaims</td> <td>number of means-plus-function claims</td> </tr> <tr> <td>meansFirst</td> <td>=1 if first claim is a means-plus-function claim</td> </tr> <tr> <td>JepsonClaims</td> <td>number of Jepson claims</td> </tr> <tr> <td>JepsonFirst</td> <td>=1 if first claim is a Jepson claim</td> </tr> </tbody> </table> <p><br>Note: The following variables/fields are currently empty (March 30, 2020); we will populate these variables/fields with data version 3.4.0.</p> <p>preambleTerm<br>preambleTermAlt<br>preambleTextStub<br>bodyLinesStep<br>bodyLinesElement<br>bodyLinesTotal</p> <p>Note: We will release the data for patents issued in 2021 with data version 3.4.0.</p> <p><br><strong>4. Coming Soon!</strong></p> <p>We are working on a number of extensions of the patccat data.</p> <p>- With data version 3.4.0, we plan to release data for all published U.S. patent applications (2001 through 2021)<br>- In late spring/early summer 2022, we will release data for patents issued by the European Patent Office (EPO) [<strong>Update: March 28, 2023</strong>: see <a href="https://doi.org/10.5281/zenodo.7776092">https://doi.org/10.5281/zenodo.7776092</a>]<br>- In late spring/early summer 2022, we will release data for patents issued by the Canadian Intellectual Property Office (CIPO)</p> <p> </p>
TecKnoGraph: Knowledge Graph from patents in C4ISTAR
<p><img src="https://github.com/nicolamelluso/TecKnoGraph-demo/blob/main/TecKnoGraph-Example%20Graph.png" alt="TecKnoGraph"></p> <p>This dataset contains a sample of Knowledge Graph (KG) created with TecKnoGraph.</p> <p>There are two files:</p> <p><strong>- TecKnoGraph-C4ISTAR-sample.csv</strong>: this file contains the KG in the form of triples where each element of the triple (source, relation, target) is tagged with categories.</p> <p><strong>- patents.zip:</strong> this file contains data about 10,000 patents; each patent corresponds to a txt file.</p> <p>There is available a demo for using TecKnoGraph from examples of input text:<br> https://nicolamelluso-tecknograph-demo-tecknograph-streamlit-cnq4sm.streamlitapp.com/</p> <p>In this repository it is possible to find also the appendix of the corresponding paper.</p>
Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis
<p>Dataset of the research paper: <strong>Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis</strong></p> <p>Existing work on the practical impact of software engineering (SE) research examines industrial relevance rather than adoption of study results, hence the question of how results have been practically applied remains open. To answer this and investigate the outcomes of impactful research, we performed a quantitative and qualitative analysis of 4 354 SE patents citing 1 690 SE papers published in four leading SE venues between 1975–2017. Moreover, we conducted a survey on 475 authors of 593 top-cited and awarded publications, achieving 26% response rate. Overall, researchers have equipped practitioners with various tools, processes, and methods, and improved many existing products. SE practice values knowledge-seeking research and is impacted by diverse cross-disciplinary SE areas. Practitioner-oriented publication venues appear more impactful than researcher-oriented ones, while industry-related tracks in conferences could enhance their impact. Some research works did not reach a wide footprint due to limited funding resources or unfavorable cost-benefit trade-off of the proposed solutions. The need for higher SE research funding could be corroborated through a dedicated empirical study. In general, the assessment of impact is subject to its definition. Therefore, academia and industry could jointly agree on a formal description to set a common ground for subsequent research on the topic.</p> <p>The following data files are included.</p> <ul> <li><em>./fields</em>: <ul> <li><strong>engi-fields.csv</strong>: Publication and PhD dissertation counts of main engineering branches</li> <li><strong>engi-fields-queries.txt</strong>: Queries applied to Elsevier's Scopus and Open Access Theses and Dissertations databases to retrieve the publication and dissertation counts</li> </ul> </li> <li><em>./patents</em>: <ul> <li><strong>sample-se-references-verified.csv</strong>: Manual verification of a random sample of references by software engineering (SE) patents to SE papers</li> <li><strong>se-cpc.tsv</strong>: Manually-identified SE-related Cooperative Patent Classification (CPC) categories</li> <li><strong>se-references-in-patents.csv</strong>: SE references made by SE patents to SE papers</li> <li><em>./patents/litigation</em>: <ul> <li><strong>case-values.csv</strong>: Manually-retrieved litigation damages of citing SE patents</li> <li><strong>lit-per-paper.csv</strong>: Litigation cases of citing SE patents</li> </ul> </li> <li><em>./patents/maintenance</em>: <ul> <li><strong>maint-code-fee-mapping.csv</strong>: Mapping of patent maintenance fee codes to their fee values</li> <li><strong>maint-fees.csv</strong>: Fee values of maintenance fee codes</li> <li><strong>maint-per-paper.csv</strong>: Maintenance fee events of citing SE patents</li> </ul> </li> <li><em>./patents/reports</em>: <ul> <li><strong>lit-sum-per-paper.csv</strong>: Counts and total damages of litigation cases of patent-cited SE papers</li> <li><strong>maint-sum-per-paper.csv</strong>: Counts and total values of maintenance fee events of patent-cited SE papers</li> <li><strong>patent-ref-counts.csv</strong>: SE patent citation counts of patent-cited SE papers</li> </ul> </li> </ul> </li> <li><em>./survey</em>: <ul> <li><strong>emse-top.csv</strong>: Most-cited papers of the Empirical Software Engineering (EMSE) journal</li> <li><strong>icse-bp.csv</strong>: Distinguished papers of the International Conference of Software Engineering (ICSE)</li> <li><strong>icse-mip.csv</strong>: Most influential ICSE papers</li> <li><strong>icse-top.csv</strong>: Most-cited ICSE papers</li> <li><strong>survey-questionnaire-emse.pdf</strong>: The EMSE survey questionnaire</li> <li><strong>survey-questionnaire.pdf</strong>: The ICSE, TSE, and TOSEM survey questionnaire</li> <li><strong>survey-responses.csv</strong>: The anonymized survey responses</li> <li><strong>tosem-top.csv</strong>: Most-cited papers of the ACM Transactions on Software Engineering and Methodology (TOSEM)</li> <li><strong>tse-top.csv</strong>: Most-cited papers of the IEEE Transactions on Software Engineering (TSE)</li> <li><em>./survey/manual-coding</em>: <ul> <li><strong>feedback.txt</strong>: Manual coding of survey feedback</li> <li><strong>practical-impact.csv</strong>: Manual coding of responses about practical impact of work</li> <li><strong>practical-impact-lack.csv</strong>: Manual coding of responses about lack of practical impact</li> <li><strong>research-methods.csv</strong>: Manual coding of additional research methods of surveyed papers</li> <li><strong>state-of-practice.csv</strong>: Manual coding of responses about changes in state of practice</li> </ul> </li> </ul> </li> <li><em>./venues</em>: <ul> <li><strong>se-venues.csv</strong>: Top SE venues according to Google Scholar Metrics</li> <li><strong>se-venues-impact.csv</strong>: SE patent citations and patent-based impact factors of SE venues</li> <li><strong>se-venues-scopus-queries.txt</strong>: Queries applied to Scopus to retrieve the publication counts of the SE venues</li> </ul> </li> </ul>
A Patent Survey on the Biotechnological Production of 2,5-Furandicarboxylic acid (FDCA): Current Trends and Challenges
<p>The production of 2,5-furandicarboxylic acid (FDCA) as a biobased commodity chemical has gained great importance over the last years. The possibility of replacing conventional polyethylene terephthalate (PET)-based plastics by polyethylene furanoate (PEF) is accelerating applied research in the direction of novel and highly productive synthesis routes to FDCA. This paper explores the patent activities related to FDCA production, with particular emphasis on the potential role that enzymatic catalytic methods may have. An increasing number of patent applications (and granted ones) have been disclosed over the last decade. The innovation is set on the development of multi-step catalytic processes, involving multi-enzymatic and chemoenzymatic cascades, or using whole-cells, either as resting (non-growing) biocatalysts or as living organisms in fermentative processes. Moreover, other innovative paths use substrates different from fructose and 5-hydroxymethylfurfural (HMF), such as gluconic acid, or furfural. Overall, opportunities for innovation exist, albeit current production metrics in biotechnological methods remain mostly at the proof-of-concept level. Large development efforts to reach industrial targets are needed, which may be stimulated by the less severe processing conditions expected for biotechnology, leading to energy savings, less by-product formation, and to the valorization of crude effluents.</p>
MADIA_732678_INN_patents_02
<p>Collection of .csv files with raw patent data, analysed to obtain landscaping information on technologies overlapping with the project. Used to produce the deliverable “Patent and scientific literature study M24” D7.6 Patent landscaping. The data where obtained on <a href="http://www.thelens.org">www.thelens.org</a></p> <p>The queries used to obtain the data were:</p> <p>A</p> <p>title:(neurodegenerative early diagnosis) OR abstract:(neurodegenerative early diagnosis) OR claims:(neurodegenerative early diagnosis) AND classification_ipcr:((G01N*) OR (H01L*))</p> <p>B</p> <p>title:(functionalized magnetic nano*) OR abstract:(functionalized magnetic nano*) OR claims: (functionalized magnetic nano*) AND classification_ipcr:((G01N*) OR (B82Y*))</p> <p>C</p> <p>title:((microfluidic) AND (chamber OR apparatus OR device)) OR abstract:((microfluidic) AND (chamber OR apparatus OR device)) OR claims:((microfluidic) AND (chamber OR apparatus OR device)) AND classification_ipcr:((B01J*) OR (B81B*))</p> <p>D</p> <p>title:((magnetic nano*) AND (sensor OR detector)) OR abstract:((magnetic nano*) AND (sensor OR detector)) OR claims: ((magnetic nano*) AND (sensor OR detector)) AND classification_ipcr:((G01R*) OR (G11B*) OR (H01L*) OR (H01F*))</p>
3PFL: Database of Patents and Publications with a Public-Funding Linkage
<p>The 3PFL database links information on patented inventions and scientific publications related to a public procurement contract or a research grant awarded by the U.S. Federal Government to detailed contract-level/grant-level information (e.g., awarding agency, recipient organization, award size). We have combined data from multiple sources, including (but not limited to) the United States Patent and Trademark Office bulk database, the Federal Procurement Database System, the Award Submission Portal (ASP), and the European Patent Office's PATSTAT database. We also provide a link to the scientific publications associated with these patents. The 3PFL database provides rich and original information that opens the door to novel empirical research in the economics of innovation and science. </p>
AI-related patents (WIPO, category G06N) and market capitalisation by companies registering at least 2 new ones in 2019, sorted into four global regions (China, USA, EEA, rest of the world)
<p>NOTE: for some reason the pptx and previews keep getting munged on this supposedly permanent arxiv, but the data is still there, unchanged, and you can see how the pptx should look in either the jpg, or the the article.</p> <p>Datasets and presentations concerning the strength of the EU and "the rest of the world" relative to China and the USA, for the purpose of illustrating and counteracting / better informing narratives concerning a "new AI cold war". The materials authored by us may be freely used under the terms of the MIT License, which appears in its entirety in both the dataset and the presentation. The other materials are only curated by us, taken from Twitter as examples of misinformation pertaining to this concern.</p> <p>As of 28 June, this work now also appears in a formal publication: Joanna J. Bryson, Helena Malikova; Is There an AI Cold War?. <em><em>Global Perspectives</em></em> 2021; 2 (1): 24803. doi: <a href="https://doi.org/10.1525/gp.2021.24803">https://doi.org/10.1525/gp.2021.24803</a></p> <p>Authors: The original analysis was conducted primarily by Malikova in collaboration with Bryson. An associated publication is anticipated where Bryson is the lead author.</p> <p>Contributors: independently followed Malikova's procedures to check her work. Inconsistencies were triple checked and resolved.</p>
The patccat classifier for patent claims - EPO edition
<p>!!! This is the <strong>EPO/European version</strong> of the patccat classifier of patent claims. !!!</p> <p>Note: We use the same approach that we use for USPTO patents. For a detailed description, see <a href="https://doi.org/10.5281/zenodo.6395307">https://doi.org/10.5281/zenodo.6395307</a>.</p> <p><strong>Data version: 3.4.0</strong></p> <p>Authors:<br> Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br> W. Keith Robinson (Wake Forest University, School of Law)<br> Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): "The Rise of Process Claims: Evidence from a Century of U.S. Patents," unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p>
Patent text: code, data, and new measures
<p>This Zenodo page describes data collection, processing, and different open access data files related to the text of USPTO patent documents. The document "Data Description Zenodo.pdf" provides more details. If you use the code or data, please cite the following paper:</p> <p>Arts S, Hou J, Gomez JC (2021). Natural language processing to identify the creation and impact of new technologies in patent text: Code, data, and new measures. <em>Research Policy</em>, 50(2), 104144. (<a href="https://doi.org/10.1016/j.respol.2020.104144">https://doi.org/10.1016/j.respol.2020.104144</a>)</p>
PatCit: A Comprehensive Dataset of Patent Citations
<p><em><strong>patCit: A Comprehensive Dataset of Patent Citations</strong></em> [<a href="https://tinyletter.com/patcit">Newsletter</a>, <a href="https://github.com/cverluise/PatCit">GitHub</a>]</p> <p>Patents are at the crossroads of many innovation nodes: science, industry, products, competition, etc. Such interactions can be identified through citations <em>in a broad sense</em>.</p> <p>It is now common to use front-page patent citations to study some aspects of the innovation system. However, <strong>there is much more buried in the Non Patent Literature (NPL) citations and in the patent text itself</strong>. <strong>patCit extracts and structures these citations.</strong></p> <blockquote> <p>Want to know more? Read patCit <a href="https://docs.google.com/presentation/d/11COlz64EZn8PipXvnDBBZI_bnDD0fpm6tyx1_EqD6lU/edit?usp=sharing">academic presentation</a> or dive into usage and technical guides on patCit <a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p> </blockquote> <p><strong><em>IN PRACTICE</em></strong></p> <p>At patCit, we are building a <em>comprehensive</em> dataset of patent citations to help the community explore this <em>terra incognita</em>. patCit has the following features:</p> <ul> <li>global coverage</li> <li>front-page and in-text citations</li> <li>all categories of NPL documents</li> </ul> <p><strong><em>Front-page</em></strong></p> <p>patCit builds on <a href="https://www.epo.org/searching-for-patents/data/bulk-data-sets/docdb.html#tab-1">DOCDB</a>, the largest database of Non Patent Literature (NPL) citations. First, we deduplicate this corpus and organize it into 10 categories (bibliographical reference, database, norm & standard, etc). Then, we design and apply category specific information extraction models using <a href="https://github.com/explosion/spaCy">spaCy</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p><strong><em>In-text</em></strong></p> <p>patCit builds on Google Patents corpus of <a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&p=patents-public-data&d=patents&t=publications&page=table">USPTO full-text patents</a>. First, we extract patent and bibliographical reference citations. Then, we parse detected in-text citations into a series of category dependent attributes using <a href="https://github.com/kermitt2/grobid">grobid</a>. Patent citations are matched with a standard publication number using the Google Patents <a href="https://patents.google.com/api/match">matching API</a> and bibliographical references are matched with a DOI using <a href="https://github.com/kermitt2/biblio-glutton">biblio-glutton</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p> </p> <p><strong>FAIR</strong></p> <p><strong>Find</strong> - The patCit dataset is available on <a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&p=patcit-public-data&page=project">BigQuery</a> in an interactive environment. For those who have a smattering of SQL, this is the perfect place to explore the data. It can also be downloaded on <a href="https://zenodo.org/record/3710994">Zenodo</a>.</p> <p><strong>Interoperate</strong> - Interoperability is at the core of patCit ambition. We take care to extract unique identifiers whenever it is possible to enable data enrichment for domain specific high quality databases. This includes the DOI, PMID and PMCID for bibliographical references, the Technical Doc Number for standards, the Accession Number for Genetic databases, the publication number for PATSTAT and Claims, etc. See specific table for more details.</p> <p><strong>Reproduce</strong> - Our <a href="https://github.com/cverluise/PatCit">gitHub</a> repository is the project factory. You can learn more about data recipes and models on the patCit <a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p>
USPTO patent data: 250k random sample and NPEs' patents
<p>This package includes:</p> <ol> <li>Two Stata .dta files consisting of information on patents assigned by the United States Patent and Trademark Office between 1976 and 2014: a random sample of 250 000 US patents, and data on patent owned by Intellectual Ventures, RPX, and several other companies. The variables for example include: grant date, application date, forward and backward citations, renewals, claims and others.</li> <li>Source codes and methods used in generating and analyzing the two data files.</li> </ol> <p>A bachelor thesis with further information will be linked here.</p>
Link Compustat – USPTO Patent Assignment Dataset
<p>This page provides the data resulting from linking assignees and assignors in the USPTO Patent Assignment Dataset to Compustat gvkeys. We work with a version of the USPTO PAD that was gracefully shared with us by Stuart Graham. Such version precedes by one year the first release available at the USPTO website (https://www.uspto.gov/ip-policy/economic-research/research-datasets/patent-assignment-dataset). The version that we use covers 5,534,135 transactions recorded at the USPTO between January 1970 and January 2013 (inclusive). While the first transaction date is January 1970, the number of transactions recorded in the initial years is negligible. Data coverage seems sufficient for the years 1981-2012.</p> <p> </p> <p>If you use the code or data, please cite the following two papers:</p> <p> </p> <p>Arque-Castells, P., and Spulber, D. (2022). Measuring the Private and Social Returns to R&D: Unintended Spillovers versus Technology Markets. Journal of Political Economy. <a href="https://doi.org/10.1086/719908">https://doi.org/10.1086/719908</a></p> <p> </p> <p>Arqué Castells, Pere and Spulber, Daniel F., Firm Matching in the Market for Technology: Business Stealing and Business Creation (September 17, 2021). Northwestern Law & Econ Research Paper No. 18-14, Available at SSRN: https://ssrn.com/abstract=3041558 or <a href="http://dx.doi.org/10.2139/ssrn.3041558">http://dx.doi.org/10.2139/ssrn.3041558</a></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.