Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.7.1
Dataset results
6 results for “entity resolution”
Figure 5. Final Ranking-Data Conflict Resolution among Same Entities in Web of Data
<p>Finally, the last score with regard to the equation 4 has been calculated and is shown in<br> figure 5. As it shows Geonames gains the highest score. Considering the proposed idea, Geonames'<br> data is the most accurate. So when we are faced with data conflicts, Geonames' data must be chosen.<br> The WordFactbook is also illustrated in the above table. We can see if WordFactbook is not<br> removed from the synonyms list it would have more accurate data compared with DBpedia data.</p>
Figure 3. Result of first fraction-Data Conflict Resolution among Same Entities in Web of Data
<p>Finally we count the entities that contain any words in the synonyms list. The result of first<br> the fraction in equation 4 is shown in figure 3.</p>
Figure 2. Ranking datasets by Page Rank-Data Conflict Resolution among Same Entities in Web of Data
<p>As depicted in Figure 2 DBpedia is the top ranked data set while GeoLinked Data and<br> Eurostat are the low ranked data sets. In this stage the data sets whose rank scores are less than half<br> of the rank scores belonging to the top ranked data set are removed from the assessment.</p>
Figure 1. General Workflow of Algorithm-.Data Conflict Resolution among Same Entities in Web of Data
<p>The Page Rank algorithm which is widely used in most search engines such as Google could<br> be easily used to rank linked data. By starting from a point and random surfing, this algorithm<br> evaluates the probability of finding any given page. The algorithm assumes a link between a page i<br> to a page j demonstrates the importance of page j. In addition, the importance of page j is associated<br> to the importance of page i itself and inversely proportional to the number of pages i point to. To<br> adapt this algorithm to web of data, any page considered as a dataset and links between pages<br> considered as links between datasets.</p>
Figure 4. Result of Second Fraction-Data Conflict Resolution among Same Entities in Web of Data
<p>In the last stage the size of the data set is examined. The number of instances of specialized<br> entities and total entities are calculated. For example "London" is an instance of a specialized entity.<br> In table 2 one of the properties of "London" is displayed in form of a triple. As is shown, the object<br> of the triple is dbpedia-owl:Place that contains "place" as a member of the specialized entity. And 'A<br> Trip to the Moon' is an instance of non-specialized entity because neither the subject nor the object<br> is in the specialized entities list. After calculating the number of specialized instances and total<br> instances the result of the second fraction of equation 4 is shown in figure 4.</p>
Datasets for Supervised Matching in Clean-Clean Entity Resolution
<p>The repository includes 13 established datasets for evaluating ML- and DL-based matching algorithms:</p> <ol> <li>Structured DBLP-ACM</li> <li>Structured DLBLP-Scholar</li> <li>Structured iTunes-Amazon</li> <li>Structured Walmart-Amazon</li> <li>Structured BeerAdvo-RateBeer</li> <li>Structured Amazon-Google Products</li> <li>Strucutred Fodors-Zagats</li> <li>Dirty DBLP-ACM</li> <li>Dirty DBLP-Scholar</li> <li>Dirty iTunes-Amazon</li> <li>Dirty Walmart-Amazon</li> <li>Textual Abt-Buy</li> <li>Textual CompanyA-CompanyB</li> </ol> <p>Additionally, the repository includes five new benchmark datasets that are drawn from the following databases using a principled approach based on DeepBlocker:</p> <ol> <li>Abt-Buy</li> <li>Amazon-Google Products</li> <li>DBLP-ACM</li> <li>IMDB-TMDB</li> <li>IMDB-TVDB</li> <li>TMDB-TVDB</li> <li>Walmart-Amazon</li> <li>DBLP-Google Scholar</li> </ol> <p>The datasets are available in different formats so that they can be processed by the following matching algorithms:</p> <ol> <li>EMTransformer</li> <li>GNEM</li> <li>HierMatcher</li> <li>Magellan</li> <li>ZeroER</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.