Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3 results for “patent claims”

Learn how ShareScore rates datasets ↗
zenodo44/100

patccat: A classifier for patent claims

<p><strong>Data version: 3.3.0</strong></p> <p>Authors:<br>Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br>W. Keith Robinson (Wake Forest University, School of Law)<br>Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p><br>1. Notes on Data Construction<br>2. Citation and Code<br>3. Description of the Data Files<br>3.1. File List<br>3.2. List of Variables for Files with Claim-Level Information<br>3.3. List of Variables for Files with Patent-Level Information<br>4. Coming Soon!</p> <p><br><strong>1. Notes on Data Construction</strong></p> <p>This is version 3.3.0 of the patccat data (patent claim classification by algorithmic text analysis).</p> <p>Patent claims define an invention. A patent application is required to have one or more claims that distinctly claim the subject matter which the patent applicant regards as her invention or discovery. We construct a classifier of patent claims that identifies three distinct claim types: process claims, product claims, and product-by-process claims.</p> <p>For this classification, we combine information obtained from both the preamble and the body of a claim. The preamble is a general description of the invention (e.g., a method, an apparatus, or a device), whereas the body identifies steps and elements (specifying in detail the invention laid out in the preamble) that the applicant is claiming as the invention. The combination of the preamble type and the body type provides us with a more detailed and more accurate classification of claims than other approaches in the literature. This approach also accounts for unconventional drafting approaches. We eventually validate our classification using close to 10,000 manually classified claims.</p> <p>The data files contain the results of our classification. We provide claim-level information for each independent claim of U.S. utility patents granted between 1836 and 2020. We also provide patent-level information, i.e., the counts of different claim types for a given patent.</p> <p>For a detailed description of our classification approach, please take a look at the accompanying paper (Ganglmair, Robinson, and Seeligson 2022).</p> <p><strong>2. Citation</strong></p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): "The Rise of Process Claims: Evidence from a Century of U.S. Patents," unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p> <p>In the paper, we document the use of process claims in the U.S. over the last century, using the patccat data. We show an increase in the annual share of process claims of about 25 percentage points (from below 10% in 1920). This rise in process intensity of patents is not limited to a few patent classes, but we observe it across a broad spectrum of technologies. Process intensity varies by applicant type: companies file more process-intense patents than individuals, and U.S. applicants file more process-intense patents than foreign applicants. We further show that patents with higher process intensity are more valuable but are not necessarily cited more often. Last, process claims are on average shorter than product claims (with the gap narrowing since the 1970s).</p> <p>We would love to see how other researchers use the data and eventually learn from it. If you have a discussion paper or a publication in which you use the data, please send us a copy at patccat.data@gmail.com.</p> <p>We will the R code used to construct the data on Github with the next data version (version 3.4.0). Contact us at b.ganglmair@gmail.com if you would like to take a look at an earlier version of the code.</p> <p><br><strong>3. Description of the Data Files</strong></p> <p>The data files contain claim-level information for independent claims of 10,140,848 U.S. utility patents granted between 1836 and 2020. The files further contain patent-level information for U.S. utility patents.</p> <p><em>3.1. File List</em></p> File list <table><tbody> <tr> <td>claims-patccat-v3-3-sample.csv</td> <td>claim-level information for independent claims of a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>claims-patccat-v3-3-1836-1919.csv</td> <td>claim-level information for independent claims of 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>claims-patccat-v3-3-1920-2020.csv</td> <td>claim-level information for independent claims of 9,102,807 patents issued between 1920 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-sample.csv</td> <td>patent-level information for a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-1836-1919.csv</td> <td>patent-level information for 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>patents-patccat-v3-3-1920-2020.csv</td> <td>patent-level information for 9,102,807 patents issued between 1920 and 2020</td> </tr> </tbody> </table> <p><br><em>3.2. List of Variables for Files with Claim-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Claim-Level Information) <table><tbody> <tr> <td>PatentClaim</td> <td>patent claim identifier; 8-digit patent number and 4-digit claim number (Ex: 01234567-0001)</td> </tr> <tr> <td>singleLine</td> <td>=1 if claim is published in single-line format</td> </tr> <tr> <td>singleReformat</td> <td>outcome code of reformating of single-line claims</td> </tr> <tr> <td>Jepson</td> <td>=1 if claim is a Jepson claim</td> </tr> <tr> <td>JepsonReformat</td> <td>outcome code of reformating of Jepson claims</td> </tr> <tr> <td>inBegin</td> <td>=1 if claim begins with the word "in"</td> </tr> <tr> <td>wordsPreamble</td> <td>number of words in the claim preamble</td> </tr> <tr> <td>wordsBody</td> <td>number of words in the claim body</td> </tr> <tr> <td>dependentClaims</td> <td>number of dependent claims that refer to this independent claim</td> </tr> <tr> <td>isMeansPreamble</td> <td>=1 if term "means" is used in the preamble</td> </tr> <tr> <td>isMeansBody</td> <td>=1 if term "means" is used in the body</td> </tr> <tr> <td>isMeans</td> <td>=1 if term "means" is used anywhere in the claim (~ means-plus-function claim)</td> </tr> <tr> <td>processPreamble</td> <td>=1 if terms "method" or "process" are used in the preamble</td> </tr> <tr> <td>processBody</td> <td>=1 if terms "method" or "process" are used in the body</td> </tr> <tr> <td>processSimple</td> <td>=1 if terms "method" or "process" are used anywhere in the claim (for simple approach of process claim classification)</td> </tr> <tr> <td>claimType</td> <td>claim type of full classification (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>preambleType</td> <td>preamble type</td> </tr> <tr> <td>preambleTerm</td> <td>keyword used to classify preamble type</td> </tr> <tr> <td>preambleTermAlt</td> <td>alternative keyword (if preambleTerm were not used)</td> </tr> <tr> <td>preambleTextStub</td> <td>first 15 words of the preamble</td> </tr> <tr> <td>bodyType</td> <td>body type</td> </tr> <tr> <td>bodyLinesStep</td> <td>number of steps in the body</td> </tr> <tr> <td>bodyLinesElement</td> <td>number of elements in the body</td> </tr> <tr> <td>bodyLinesTotal</td> <td>total number of identified lines in the body</td> </tr> <tr> <td>label</td> <td>2-character label of the preamble-body combination; classification table maps label to claim type</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><em>3.3. List of Variables for Files with Patent-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Patent-Level Information) <table><tbody> <tr> <td>patent_id</td> <td>U.S. patent number (8-digit patent number)</td> </tr> <tr> <td>claims</td> <td>number of independent claims (the sum of the four claim types: 0, 1, 2, and 3)</td> </tr> <tr> <td>noCategory</td> <td>number of claims without a classified type</td> </tr> <tr> <td>processClaims</td> <td>number of process claims</td> </tr> <tr> <td>productClaims</td> <td>number of product claims</td> </tr> <tr> <td>prodByProcessClaims</td> <td>number of product-by-process claims</td> </tr> <tr> <td>firstClaim</td> <td>type of the first claim (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>simpleProcessClaims</td> <td>number of process claims by simple approach (terms "method" or "process" anywhere in the claim)</td> </tr> <tr> <td>simpleProcessPreamble</td> <td>number of process claims by simple approach (terms "method" or "process" in the preamble)</td> </tr> <tr> <td>meansClaims</td> <td>number of means-plus-function claims</td> </tr> <tr> <td>meansFirst</td> <td>=1 if first claim is a means-plus-function claim</td> </tr> <tr> <td>JepsonClaims</td> <td>number of Jepson claims</td> </tr> <tr> <td>JepsonFirst</td> <td>=1 if first claim is a Jepson claim</td> </tr> </tbody> </table> <p><br>Note: The following variables/fields are currently empty (March 30, 2020); we will populate these variables/fields with data version 3.4.0.</p> <p>preambleTerm<br>preambleTermAlt<br>preambleTextStub<br>bodyLinesStep<br>bodyLinesElement<br>bodyLinesTotal</p> <p>Note: We will release the data for patents issued in 2021 with data version 3.4.0.</p> <p><br><strong>4. Coming Soon!</strong></p> <p>We are working on a number of extensions of the patccat data.</p> <p>- With data version 3.4.0, we plan to release data for all published U.S. patent applications (2001 through 2021)<br>- In late spring/early summer 2022, we will release data for patents issued by the European Patent Office (EPO) [<strong>Update: March 28, 2023</strong>: see <a href="https://doi.org/10.5281/zenodo.7776092">https://doi.org/10.5281/zenodo.7776092</a>]<br>- In late spring/early summer 2022, we will release data for patents issued by the Canadian Intellectual Property Office (CIPO)</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

The patccat classifier for patent claims - EPO edition

<p>!!! This is the <strong>EPO/European version</strong> of the patccat classifier of patent claims. !!!</p> <p>Note: We use the same approach that we use for USPTO patents. For a detailed description, see <a href="https://doi.org/10.5281/zenodo.6395307">https://doi.org/10.5281/zenodo.6395307</a>.</p> <p><strong>Data version: 3.4.0</strong></p> <p>Authors:<br> Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br> W. Keith Robinson (Wake Forest University, School of Law)<br> Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): &quot;The Rise of Process Claims: Evidence from a Century of U.S. Patents,&quot; unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo24/100

Biological sequences named and claimed in US patents and patent applications, CAMBIA Patent Lens Archived 2013-07-27

<p>Archived PatentLens sequence data released under Attribution-NonCommercial-ShareAlike 2.5 Generic (CC BY-NC-SA 2.5)</p> <p>http://web.archive.org/web/20130410080554/http://www.patentlens.net/<br> http://web.archive.org/web/20130414210513/http://www.patentlens.net/sequence/<br> http://web.archive.org/web/20130727112746/http://www.patentlens.net/sequence/sw/<br> http://web.archive.org/web/20130727125050/http://www.patentlens.net/sequence/US_A/<br> http://web.archive.org/web/20130412104754/http://www.patentlens.net/sequence/US_B/</p>

opencc-by-nc-sa-2.5Jul 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record