Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,644
datasets available to search
ShareScore release 0.9.0
Dataset results
1,644 results for “italian”
Investigating the effects of COVID‑19 lockdown on Italian children and adolescents with and without neurodevelopmental disorders: a cross‑sectional study - DATASET
<p>Dataset to support the findings in the journal paper titled "Investigating the effects of COVID‑19 lockdown on Italian children and adolescents with and without neurodevelopmental disorders: a cross‑sectional study".</p> <p>Each row is a different subject.</p> <p>Each column represents an answer to the questionnaire. For single choice questions, the answer was reported as-is (Italian). For multiple choice questions, the alternatives where splitted in several columns and the answer was coded as 0/1 (one hot encoding). For the "school" column, 2=primary school, 3=middle school, 4=high school. For the "school.class" column, classes from 4 to 8 belong to primary school, from first to fifth grade; classes from 9 to 11 belong to middle school, from first to third grade; classes from 12 to 16 belong to high school, from first to fifth grade.</p>
3D Adjoint tomography model of the Italian lithosphere
<p>The project IMAGINE_IT (PI Dr. Dimitri Komatitsch) received 40 million CPU-hours on the Tier-0 GENCI/TGCC CURIE supercomputer as a winner of the 9<sup>th</sup> PRACE consortium call (2014). </p> <p>The awarded computational resources allowed us to construct a new 3D tomographic model for the Italian lithosphere, <em>Im25</em>,<em> </em>by combining spectral-element three-dimensional wavefield simulations and an adjoint-state method.</p> <p>To obtain the final model <em>Im25, </em>we performed 25 adjoint tomography iterations and two source inversion iterations for the initial wavespeed model and an intermediate one.</p> <p>The 3D model <em>Im25</em> resolves P- and S-wavespeed values in the Italian lithosphere for frequencies up to ~0.1 Hz (i.e., periods down to ~10 s). It provides improved images of the subsurface of the peninsula and its neighborhood, highlighting the complex structure of the Adriatic plate on the east of the Italian coast, the debated plumbing system of Mount Etna volcano, and the distribution of fluids and gas (CO<sub>2</sub>) in correlation with the Italian seismicity. </p> <p>The obtained <em>V<sub>p</sub></em> and <em>V<sub>s</sub></em><sub> </sub>models are in agreement with high resolution models developed by classical tomography studies in Italy at local scale, and the retrieved <em>V<sub>p</sub></em>/<em>V<sub>s</sub></em> values correspond to the interpretations on fluids and gas distribution in the Italian subsurface from studies in various geophysical fields. The <em>V<sub>s</sub></em> <em>Im25</em> model also compares favorably with <em>V<sub>s</sub></em> profiles extracted from other studies that use different techniques and datasets, and focus on specific zones of Italy.</p> <p>The provided dataset contains the values of <em>V</em><sub><em>p</em></sub> and <em>V<sub>s</sub></em> (in m/s) for model <em>Im25</em> at each point of the geometrical grid used to discretize the study volume.</p>
Datasets and results of the paper titled "Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"
<p>These are the <strong>input datasets</strong> and the <strong>results of the analyses</strong> reported on the paper titled <strong>"Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"</strong>.</p> <p><strong>Abstract:</strong> </p> <p>The aim of this paper is to study the role of citation network measures in the assessment of scientific maturity. Referring to the case of the Italian national scientific qualification (ASN), we investigate if there is a relationship between citation network indices and the results of the researchers’ evaluation procedures. In particular, we want to understand if network measures can enhance the prediction accuracy of the results of the evaluation procedures beyond basic performance indices. Moreover, we want to highlight which citation network indices prove to be more relevant in explaining the ASN results, and if quantitative indices used in the citation-based disciplines assessment can replace the citation network measures in non-citation-based disciplines. Data concerning Statistics and Computer Science disciplines are collected from different sources (ASN, Italian Ministry of University and Research, and Scopus) and processed in order to calculate the citation-based measures used in this study. Following, we apply classification models to estimate the effects of network variables. We find that network measures are strongly related to the results of the ASN and significantly improve the explanatory power of the models, especially for the research fields of Statistics. Additionally, citation networks in the specific sub-disciplines are far more relevant than those in the general disciplines. Finally, results show that the citation network measures are not a substitute of the citation-based bibliometric indices.</p> <p><strong>Code</strong></p> <p>The code to collect and process the data used in this paper is available on GitHub at <a href="https://github.com/DigitalDataLab/ASN16-18_CitationNetwork">https://github.com/DigitalDataLab/ASN16-18_CitationNetwork</a><strong>.</strong> </p> <p><strong>Dataset description</strong></p> <p>The files <strong>AdjacencyMatrix_01B1.csv</strong>, <strong>AdjacencyMatrix_09H1.csv</strong>, <strong>AdjacencyMatrix_13D1.csv</strong>, <strong>AdjacencyMatrix_13D2.csv</strong> and <strong>AdjacencyMatrix_13D3.csv</strong> are the citation matrices for Italian academics (i.e. ASN candidates and permanent positions in the Italian academic system) in the Recruitment Fields (RFs) 01/B1, 09/H1, 13/D1, 13/D2 and 13/D3, respectively.</p> <p>The files <strong>AdjacencyMatrix_CS.csv</strong> and <strong>AdjacencyMatrix_ST.csv</strong> are the citation matrices for the Italian academics in the Computer Science disciplines (i.e. RFs 01/B1 and 09/H1) and the Statistical disciplines (i.e. RFs 13/D1, 13/D2 and 13/D3), respectively.</p> <p>The files <strong>CS_01B1_1.csv, CS_09H1_1.csv, ST_13D1_1.csv, ST_13D2_1.csv</strong> and <strong>ST_13D3_1.csv</strong> contain the data used to build the logistic regression models presented in the paper for the Italian academics at the Full Professor (FP) level.</p> <p>The files <strong>CS_01B1_2.csv, CS_09H1_2.csv, ST_13D1_2.csv, ST_13D2_2.csv</strong> and <strong>ST_13D3_2.csv</strong> contain the data used to build the logistic regression models presented in the paper for the Italian academics at the Associate Professor (AP) level.</p> <p>The file <strong>Codebook.pdf</strong> is the codebook of the previous ten files.</p> <p>The file <strong>Appendix.pdf</strong> contains the final results of the stepwise logistic regressions computed for each level (i.e. Full Professor and Associate Professor) and Recruitment Field in the Computer Science and Statistics disciplines.</p> <p>The file <strong>NormalityAssessment.pdf</strong> contains the normality assessment of citation network indices. </p>
LADDER. Learners' digital communication: a corpus for pragmatic competences in Italian L1/L2
<p> </p> <p><strong>Ladder</strong>. A Corpus of Computer-Mediated Communication for the Analysis of the Acquisition of Pragmalinguistic Competences by German-Speaking Learners of Italian.</p> <p> </p> <p> </p> <p> </p> <p>Project description:</p> <p>Many recent research projects (Artoni, Benigni, & Nuzzo, 2020; Cortés Velásquez & Nuzzo, 2017; Nuzzo & Cortés Velásquez, 2020) have underlined the usefulness of creating and analyzing corpora for teaching pragmatics, which, unlike other linguistic levels such as syntax, cannot be explained by rules but only by reference to tendential values or more or less appropriate choices in a given context. This is even more true for interactions via digital media, such as email and instant-messaging services, which have little place in manuals or L2 courses and for which learners have few reference models (Brocca, 2021; Trubnikova & Garofolin, 2020).</p> <p>Data collection:</p> <p>Data were collected from April 2020 to April 2021 with the help of a discourse completion task (DCT). The data consists of emails and instant messages. The informants are (i) German learners of Italian between A2-C1 level according to the CEFR and most of them are students living in Tyrol (Austria) and (ii) native speakers of Italian most of whom are students from Rome (Italy). The data of the learners were collected by students of the undergraduate seminar “Insegnare la pragmatica” which is part of the compulsory module 2b for student teachers at the Institute of Didactics of the University of Innsbruck. The data of the native speakers were collected in large part from students in foreign languages at the University RomaTre thanks to the collaboration with Prof. Elena Nuzzo.</p> <p>The DCTs have been conducted with online questionnaires. Along with the texts, metadata were also registered with the help of an online questionnaire giving sociolinguistic information about the informant (age, self-assessed language level, place of residence, native language, etc.). The DCTs aim to elicit linguistic acts of request and refusal in increasing levels of social distance and different media (Taguchi & Roever, 2017, pp. 85, 231; Hinger et al. 2018: 148). The DCTs elicit different speech acts (requests and refusals) with different degrees of formality (study/work or free time), directed at different people (lecturer, friend, boss) and in different media (mail or instant messaging). The scenarios represent authentic circumstances for the students. The following table shows the situations that were studied:</p> <p> </p> <p><strong>Email</strong></p> <p>high level of social distance between sender and recipient</p> <p>Scenario 1: Sender is asking for something that he/she is not entitled to</p> <p>Scenario 2: Sender is asking for something that he/she is entitled to</p> <p><strong><em>WhatsApp</em></strong><strong> messages</strong></p> <p>a) low level of social distance between sender and recipient</p> <p>Scenario 1: Request</p> <p>Scenario 2: Rejecting a request</p> <p>Scenario 3: Short-notice cancellation of an invitation</p> <p>b) medium level of social distance between sender and recipient</p> <p>Scenario 4: Request</p> <p>Scenario 5: Rejecting a request</p> <p>Scenario 6: Short-term rejection of an invitation</p> <p> </p> <p> </p> <p>The <em>WhatsApp</em> messages, which are exemplary of the text type instant messaging, were produced directly with the cell phone. The metadata were subsequently associated with the respective messages in an Excel spreadsheet. All personal data were anonymized.</p> <p>The prompts were presented in Italian, as follows:</p> <p>Mail</p> <p><strong>Mail a)</strong> Immagina di star facendo un corso con il Dr. Nicola Brocca. Domani devi fare una presentazione in classe. Non hai avuto tempo per studiare perché dovevi prepararti a un esame di inglese e ti accorgi che il materiale da presentare è più di quello che avevi previsto. Scrivi una mail al professore: la tua speranza è spostare la presentazione.</p> <p>Engl: Imagine you are taking a course with Dr. Nicola Brocca. Tomorrow you have to give a presentation in class. You had no time to study because you had to prepare for an English exam, and you realize that there is more material to present than you had imagined. You write an email to the professor: your hope is to reschedule the presentation.</p> <p><strong>Mail b)</strong> Hai fatto un corso con il Dr. Brocca. Hai consegnato il tuo portfolio il 01.02.2020 adesso è il 01.03.2020 e non hai ancora ricevuto il voto. Ti serve il voto per registrarti per una borsa di studio. Manda una mail al prof.: il tuo obiettivo è ricevere il voto al più presto</p> <p>Engl: You have taken a course with Dr. Brocca. You turned in your portfolio on 02/01/2020, it is now 03/01/2020 and you have not received the grade yet. You need the grade to register for a scholarship. Send an email to the professor: your goal is to receive the grade as soon as possible.</p> <p> </p> <p><em>WhatsApp</em> messages</p> <p><strong>1.</strong> Sei in Erasmus in Italia. Avete creato una chat con 10 compagni di corso. Hai perso la tua tessera della biblioteca a vuoi chiedere se qualcuno ti può aiutare perché ti serve un libro entro domani...per esempio prestandoti la sua. Cosa scrivi?</p> <p>Engl: You are taking part in the Erasmus program in Italy. You have created a chat with 10 classmates. You lost your library card and want to ask if someone can help you because you need a book by tomorrow.... E.g. by lending you their card. What do you write?</p> <p><strong>2.</strong> Ricevi questo messaggio da un amico/a che fa un seminario con te: "Ciao, sono a corto di tempo. Ho visto che hai preso 30 all'esame. Potresti darmi una mano e restare con me in biblioteca oggi?" Non vuoi aiutare il tuo amico. Come reagisci?</p> <p>Engl: You receive this message from a friend who is attending a seminar with you: "Hello, I'm running out of time. I saw that you got a 30 on the exam. Could you help me and stay with me in the library today?" You don't want to help the friend. How do you respond?</p> <p><strong>3.</strong> Cinque giorni fa hai promesso ad un/a amico/a che questa sera sareste andati al cinema assieme. Però hai cambiato idea. Cosa fai? Cosa scrivi?</p> <p> Engl: Five days ago, you promised a friend that tonight you would go to the movies together. But you changed your mind. What would you do? What do you write?</p> <p> </p> <p><strong>4.</strong> Sei al lavoro e hai smarrito il documento elettronico per entrare nel parcheggio. Sei nuovo in questo gruppo di lavoro e hai solo il numero del tuo diretto superiore. Gli mandi un messaggio per chiedergli se ti può aiutare.</p> <p>Engl: You are at work and have lost your electronic badge to enter the parking lot. You are new to this work group and only have the number of your direct supervisor. You send him/her a message and ask if he/she can help you.</p> <p> </p> <p><strong>5.</strong> Ricevi questo messaggio dal/la tuo/a superiore. "Gentile collega, domani c'è una scadenza importante. Per caso sarebbe in grado di restare oggi in ufficio oltre l'orario?" Non vuoi restare in ufficio oltre il normale. Come reagisci?</p> <p>Engl: You receive this message from your supervisor. "Dear colleague, tomorrow is an important appointment. Would you be able to stay in the office after hours today?" You don't want to stay in the office beyond normal working hours. How do you respond?</p> <p> </p> <p><strong>6.</strong> Cinque giorni fa hai promesso al/la tuo/a superiore che oggi saresti andato a una cena di lavoro. Però devi disdire. Cosa fai?</p> <p>Engl: Five days ago, you promised your superior that you would go to a business dinner today. However, you have to cancel. What do you do?</p> <p> </p> <p>The corpus, which was first collected in .xlsx format, was exported to XML format and CSV format in cooperation with Joseph Wang-Kathrein (Brenner Archive Research Center). It was ensured that the emoticons and special characters were also transferred unchanged in the conversion process. These formats allow long-term archiving and significantly facilitate data exchange.</p> <p>The size of the corpus (as of May 2021, version Ladder 1.0):</p> <p>The LADDER corpus includes emails and instant-messaging messages amounting to 18,935 tokens and 33,966 tokens respectively. The corpus of <em>WhatsApp</em> messages consists of a total of 1,204 messages from 80 native speakers and 114 learners. The corpus of emails consists of a total of 235 emails from 78 native-speaker informants and 38 learners. The amount of data allows a qualitatively relevant comparison in sub-corpora e.g. language levels.</p> <p>The size of the corpus is necessarily limited quantitatively, as data collection must be done manually through individual DCT management and metadata checking. The major bottleneck is currently the annotation of socio-pragmatic aspects, a process that is difficult to automate and that needs to be conducted through cross-annotation by multiple annotators.</p> <p>Some students' works on the corpus have been collected and are accessible via the following link: https://ladder.hypotheses.org/</p> <p> </p> <p>Bibliography:</p> <p>Artoni, D., Benigni, V., & Nuzzo, E. (2020), "Pragmatic instruction in L2-Russian: a study on requests and advice" in <em>Instructed Second Language Acquisition, 4</em>(1), 62-95. doi:10.1558/isla.39864</p> <p>Brocca, N. (2021), "LADDER: La costruzione e analisi di un corpus di scritture digitali per l’insegnamento della pragmatica in L2" in <em>Italiano Lingua Due, 13</em>(1 (2021)).</p> <p>Cortés Velásquez, D., & Nuzzo, E. (2017), "Disdire un appuntamento: spunti per la didattica dell'italiano L2 a partire da un corpus di parlanti nativi" in <em>Italiano Lingua Due, 1</em>, 17-36.</p> <p>Hinger, B., Stadler, W., Schmiderer, K., Bauer, M., (Hrg.) (2018). Testen und Bewerten fremdsprachlicher Kompetenzen. Tübingen: Narr Francke Attempto Verlag.</p> <p>Nuzzo, E., & Cortés Velásquez, D. (2020), "Canceling Last Minute in Italian and Colombian Spanish: A Cross-Cultural Account of Pragmalinguistic Strategies" in <em>Corpus Pragmatics, 4</em>, 1-26. doi:10.1007/s41701-020-00084-y</p> <p>Taguchi, N., & Roever, C. (2017), <em>Second language pragmatics</em>: Oxford: Oxford University Press.</p> <p>Trubnikova, V., & Garofolin, B. (2020), <em>Lingua e interazione. Insegnare la pragmatica a scuola</em>. Pisa: ETS.</p> <p> </p>
AIDA (Archive of Italian radiocarbon DAtes)
<p>The archive <strong>AIDA</strong> provides a collation of <strong>4,629</strong> radiocarbon dates from <strong>1,050</strong> archaeological sites in Italy from the Late Mesolithic until Late Antiquity (11 - 1.5 kya BP). These dates have been collected from existing online digital archives, and electronic and print original publications. </p> <p>List of versions:</p> <ul> <li><strong>5.0</strong> 9 April 2022 — 589 new dates added (update of the files 'References.txt', 'nerd.csv', and 'Readme.md').</li> <li><strong>4.0</strong> 3 March 2022 — 35 new dates added (update of the files 'References.txt', 'nerd.csv', and 'Readme.md').</li> <li><strong>3.0</strong> 13 January 2022 — Removal of some duplicates and 4 new dates added (update of the files 'References.txt', 'nerd.csv', and 'Readme.md').</li> <li><strong>2.0</strong> 13 January 2022 — Removal of some duplicates and 4 new dates added (update of the files 'References.txt', 'nerd.csv', and 'Readme.md').</li> <li><strong>1.0</strong> 3 August 2021 — First public release of the dataset on Zenodo</li> </ul>
Corpus of the Epigraphy of the Italian Peninsula in the 1st Millennium BCE
<p>The Corpus of the Epigraphy of the Italian Peninsula in the 1st Millennium BCE, or CEIPoM, is a linguistic database focusing on the Italian peninsula in the first millennium BCE. Currently, it covers Messapic, Venetic, the Sabellic languages and epigraphic Latin up to about 100 BCE.</p>
Polifonia Corpus - Periodicals Module Metadata - Italian Language
<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Italian Soil Information System
<p>The application offers both soil and climatic information, related to 1:500,000 scale geography. The soil information systems of Italy is made up of a hierarchy of geographical layers, which includes soil regions, aimed at correlating the soils of Italy with the other European countries, soil systems, for the correlation of soil at the national level, and soil sub-systems, for the regional level. The databases provides an inventory of Italian soilscapes at two reference scale, Soil region (1:5,000,000) and Soil systems (1:500,000). Relation between entities (soil-typological-unit, soil derived profile, benchmark soil profile) and geography (soil mapping units), was also stored. Geographical layers (administrative boundaries, regionand province main towns, soil regions, and rasters of climatic variables) and correlation entities was also stored. Soil regions are a regionally restricted part of the soil cover characterized by a typical climate and parent-material association. Soil systems illustrate main Italian soilscapes and are composed of homogeneous areas as for physiography, lithology, river drainage network, and land cover. Each cartographic unit was described as combination of: i) a major landform, established according to main morphological process, slope, hypsometry, kind and degree of drainage, ii) two lithological types, iii) three land cover attributes. Maximum seven land components were recognized in each land system. A land component was a specific combination of morphology, lithology and land cover that was not delineated. 1412 soil observations have been stored on the Ms Access database mostly complete with 4284 analyzed soil horizons for the most common analytical parameters (pH in water; Carbon (C) - organic; Carbonate (CO3--) - Total; Clay, Sand, and Silt fraction; Available water capacity - estimated volumetric) and 2039 photos. Several climatic variables, relevant for soil evaluation and management, have been collected and stored. Climatical maps have been produced by spatialization of longterm statistics related to meteorological stations. The map of precipitation (AnnualRainfall) was obtained by ordinary kriging of 1,613 stations completed of long termannual average values. Mean annual air temperature was obtained by ordinary kriging of 944 stations of long-term average annual air temperature. The humidity index was obtained with a simple calculation (Annual Rainfall/Mean Annual Air Temperature). Average temperature of the soil to 50 cm was calculated on the basis of the average air temperature of the long term and field capacity of the soil at the same depth, accordingto the equation: Tm = soil to 0.5 m tm air + (field capacity to 0.5 m - 20.7) / 7.9 (Constantini et al., 1999). The mean annual soil temperatures map was obtained byordinary kriging of 6,660 data of average soil temperature at 50 cm.</p>
The Italian Agricultural Regions Dataset
<p>A georeferenced dataset in delineating the boundaries of the Italian Agricultural Regions (Figure_map.png) as defined by INEA (CREA) and used in the FADN/RICA database. This dataset, in a shapefile format, (RegAgr.zip) opens the possibility of conducting geographical and climatic analyses on thousands of data on farms sampled every year in Italy by the FADN/RICA network.</p> <p>The dataset is accompanied by supplementary information on the municipalities belonging to the Agronomic Regions (ARs). Specifically, for each AR, the list of statistical codes and names of administrative units (municipalities, provinces, and regions) used by ISTAT, the Italian National Institute of Statistics, is provided (attributes.zip). This way, users can easily trace the municipalities and related information that comprise each AR. Moreover, the dataset also includes the source and ancillary data used to build the dataset (SOURCE_DATA.zip).</p>
SIMPITIKI corpus for simplification in Italian
<p>SIMPITIKI is a Simplification corpus for Italian and it consists of two sets of simplified pairs: the first one is harvested from the Italian Wikipedia in a semi-automatic way; the second one is manually annotated sentence-by-sentence from documents in the administrative domain.</p> <p>For more details, see https://github.com/dhfbk/simpitiki</p>
Italian Lexical Simplification Benchmark
<p>The corpus is a manually created benchmark to evaluate the performance of Italian lexical simplification systems. It contains 901 pairs of complex sentences and their simplified version at the lexical level (i.e. replacement of a difficult term or phrase with a simpler synonym). The dataset and a system using the benchmark are described in the paper "The impact of phrases on Italian lexical simplification" <a href="https://zenodo.org/record/1048874">https://zenodo.org/record/1048874</a></p>
MERIDA - MEteorological Reanalysis Italian DAtaset
<p>The new <strong>ME</strong>teorological <strong>R</strong>eanalysis <strong>I</strong>talian <strong>DA</strong>taset (<strong>MERIDA</strong>) has been developed to cope with the increasing weather extremes of the last 20 years, which caused several disruptions to the Italian electric system. This work has been developed following the indications emerged from the “Resilience Working Table” set up by the Italian Regulatory Authority for Energy, Networks and the Environment (ARERA). MERIDA is able to respond to the energy stakeholders, who need reliable meteorological data to implement effective adaptation strategies to operate the electric system safely.</p> <p>MERIDA consists of a dynamical downscaling of the ERA5 global reanalysis using the mesoscale model WRF-ARW. ERA5 data are retrieved with a 3-hourly temporal resolution to assure good temporal consistency. Temperature data from the SYNOP Air Force stations are also retrieved to be ingested in the WRF simulations at 3-hourly temporal resolution.</p> <p>The computational domain of MERIDA consists of 2 grids with horizontal resolution of 21 km and 7 km respectively, with the internal grid centered over Italy.</p> <p>The meteorological fields of MERIDA are open access and distributed in NETCDF file format on a regular lat-lon grid of 0.07° resolution.</p> <p>A subset of the dataset is available here for download for the period 2000-2018. The full dataset covering the period 1990-2019, and continuoulsy updated, is available at the following website: <a href="http://merida.rse-web.it">http://merida.rse-web.it/</a> </p> <p>The following subset of meteorological fields is available here for download:</p> <ul> <li>T2 - 2m temperature (K)</li> <li>PREC - Total Precipitation (mm/h)</li> <li>U10 - 10m u-component of wind (m/s)</li> <li>V10 - 10m v-component of wind (m/s)</li> <li>PSFC - Surface pressure (Pa)</li> <li>Q2 - 2m Specific Humidity (Kg/Kg)</li> <li>SWDIR - Direct global short-wave radiation (W/m<sup>2</sup>)</li> <li>SWDIF - Diffuse global short-wave radiation (W/m<sup>2</sup>)</li> <li>MSLP - Mean Sea Level Pressure (Pa)</li> <li>SNEQV - Snow Water Equivalent (mm)</li> <li>SOIL_T - Soil Temperature - Layer 5 cm (K)</li> <li>SOIL_M - Soil Moisture - Layer 5 cm (m<sup>3</sup>/m<sup>3</sup>)</li> </ul> <p>The following variables are available under request:</p> <ul> <li>SOIL_T - Soil Temperature (K, Layers: 5,25,70,150 cm)</li> <li>SOIL_M - Soil Moisture - Layer 5 cm (m<sup>3</sup>/m<sup>3</sup>, Layers: 5,25,70,150 cm)</li> <li>TT - Temperature (K, Pressure levels: 850,700,500 hPa)</li> <li>RH - Relative Humidity (%, Pressure levels: 850,700,500 hPa)</li> <li>GHT - Geopotential Height (gpm, Pressure levels: 850,700,500 hPa)</li> <li>UU - u-component of wind (m/s, Pressure levels: 850,700,500 hPa)</li> <li>VV - v-component of wind (m/s, Pressure levels: 850,700,500 hPa)</li> <li>TG - Ground Temperature (K)</li> <li>HFX - Sensible Heat Flux (W/m<sup>2</sup>)</li> <li>LH - Latent Heat Flux (W/m<sup>2</sup>)</li> <li>GRDFLX - Ground Flux (W/m<sup>2</sup>)</li> <li>TR - Transpiration Flux (W/m<sup>2</sup>)</li> </ul> <p>All the variables not included for download may be downloaded at : <a href="http://merida.rse-web.it">http://merida.rse-web.it/</a></p> <p>or requested at:</p> <ul> <li>riccardo.bonanno@rse-web.it</li> <li>matteo.lacavalla@rse-web.it</li> <li>simone.sperati@rse-web.it</li> </ul> <p> </p>
Wikipedia: wikipedia-it (Italian)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>Chiunque può collaborare a Wikipedia, creando una nuova voce o migliorando i contenuti di quelle già esistenti, e chiunque voglia contribuire al progetto, nel pieno rispetto dei suoi Cinque pilastri e, soprattutto, utilizzando il buon senso, è sempre ben accetto. <p></p>https://it.wikipedia.org
Italian COVID-19 Integrated Surveillance Dataset (v42.0.0)
<p><strong>Abstract</strong></p> <p>COVID-19 integrated surveillance data provided by the <a href="http://www.iss.it/">Italian National Institute of Health</a> and processed via <a href="https://github.com/InPhyT/UnrollingAverages.jl">UnrollingAverages.jl</a> to deconvolve the weekly simple moving averages.</p> <p><strong>Overview</strong> </p> <p>Every week the National Institute for Nuclear Physics (<a href="https://home.infn.it/it/">INFN</a>) imports an anonymous individual-level dataset from the Italian National Institute of Health (<a href="https://www.iss.it/">ISS</a>) and converts it into an incidence time series data organized by date of event and disaggregated by sex, age and administrative level with a consolidation period of approximately two weeks. The information available to the <a href="https://home.infn.it/it/">INFN</a> is summarised in the following <a href="https://covid19.infn.it/iss/campi-iss.pdf">meta-table</a>.</p> <p><strong>Output Data </strong></p> <p>The output data has been stored <a href="https://github.com/InPhyT/COVID19-Italy-Integrated-Surveillance-Data/tree/main/3_output/data">here</a> and contain the following information:</p> <ul> <li>Reconstructed daily time series of <strong>confirmed cases by date of diagnosis</strong> stratified by sex and age at the regional level;</li> <li>Reconstructed daily time series of <strong>symptomatic cases by date of symptoms onset</strong> stratified by sex and age at the regional level;</li> <li>Reconstructed daily time series of <strong>ordinary hospital admissions</strong> by date of admission stratified by sex and age at the regional level;</li> <li>Reconstructed daily time series of <strong>intensive hospital admissions</strong> by date of admission stratified by sex and age at the regional level;</li> <li>Reconstructed daily time series of <strong>deceased cases by date of death</strong> stratified by sex and age at the regional level.</li> </ul>
ELABORATION OF THE ITALIAN PORTION OF THE GLOBAL SOIL ORGANIC CARBON MAP (GSOCMAP)
<p>The Global Soil Organic Carbon map (GSOCmap) published by the Food and Agriculture Organization<br> constitutes a baseline estimation of soil organic carbon stock (CS, ton ha–1) from 0 to 30 cm, on a grid at 30 arc-seconds<br> resolution (approximately 1 x 1 km). It has been produced for the Italian territory by the Italian Soil Partnership (ISP): a<br> national hub of institutions dealing with soils, either academic/research institutions, and regional soil services (RSS). The<br> RSS are the main soil data owners in Italy and play a central role in the elaboration of policies for soil management. The<br> RSS adhering to the ISP are: Calabria, Campania, Emilia Romagna, Friuli Venezia Giulia, Liguria, Lombardia, Marche,<br> Piemonte, Puglia, Sicilia, Toscana, and Veneto. A national soil database is maintained by the Consiglio per la Ricerca e<br> l'Analisi dell'Economia Agraria (CREA). The RSS contributed with soil data, with mean density of 1 point per 50 square<br> kilometres, selecting data analysed for soil organic carbon content (SOC, dag kg-1), which were representative and well<br> distributed for the following environmental covariates: land use, geomorphology, and climate. The data were selected inbetween<br> 1990 al 2013. This was necessary in order to exclude the effect of the new soil protection policies of the Rural<br> Development Programme 2014-2020. For the RSS not included in the ISP, the data were selected from the national soil<br> database. 6748 point data were finally selected. SOC values obtained with the Springer and Klee and flash combustion<br> elemental analyser methods were retained for elaborations, because the 2 methods, were found to give statistically<br> equivalent results. SOC values obtained with Walkey and Black method were, instead, corrected with an empirical factor<br> of 1.3. 2292 of the 6748 point data had also measured bulk density (BD, Mg m–3). Pedotransfer functions were calibrated<br> to estimate BD were measured BD were missing, with the following as auxiliary variables: land use, soil regions, texture,<br> and SOC. The carbon stock (CS, ton ha–1) was calculated by multiplying: 0.3 (m) * SOC (dag kg-1) * fine earth fraction (1 -<br> skeletal content expressed as daL m–3) * BD (Mg m–3). CS of the first 30 cm depth was calculated as depth-weighted<br> average. A spatial statistics method was used for the CS interpolation. The following auxiliary variables were used: soil<br> regions, soil subregions, Corine land cover 2006, lithology, soils affected by natural constrains (gleyic, histic, vertic,<br> coarse, shallow, arenic, sodic, and acid), sand content, silt content, 30-m aster-DEM, distance from coast, distance from<br> relieves, soil aridity index, annual mean precipitations, mean annual air temperature, soil inorganic carbon, and soil<br> depth. For the soil region of Po valley, the land units at 1:250,000 scale were also used. The interpolation method was a<br> general linear regression for the soil regions of Po valley, and a radial basis function for the remaining Italian territory.<br> The 6748 point data were divided, by spatial random sampling, into 10 subsets. Ten interpolations were produced, each<br> time leaving out 1/10 of the dataset. Average (fig. 1), standard deviation and confidence intervals of these 10<br> interpolations were calculated. Mean Absolute Errors (MAE) and Root Mean Squared Errors (RMSE) were respectively<br> 25.5 and 36.4 Mg/ha.</p> <p>A.85 Italy Map source: Country submission Point data Number of samples: 6748 Sampling period: 1990-2013 SOC analysis method: SOC values obtained with the Springer and Klee and ’flash combustion elemental analyser’ methods were retained for elaborations. Uncorrected values obtained by the Walkey and Black method were corrected with an empirical linear equation, based on previous studies and as recommended by the Italian official methods. BD analysis method: Undisturbed sampling, core method and pit method Mapping method Mapping method details: Neural Networks and GLM, according to soil region Validation statistics: Mean Error (ME) of the prediction is 1.688 Mg/ha, MAE 25.57 Mg/ha, Root Mean Squared Error (RMSE) is 36.24 Mg/ha. Contact Data Holder: Research centre for agriculture and environment Contact: CREA Consiglio per la ricerca in agricoltura e l’analisi dell’economia agraria edoardo.costantini@crea.gov.it</p>
Meteorological variables for Agriculture: a Dataset for the Italian Area (MADIA)
<p> </p> <p>The dataset is the supplementary material for the following journal paper:</p> <p>Parisse B.*, Alilla R., Pepe A.G., De Natale F., <em>MADIA - Meteorological variables for Agriculture: a Dataset for the Italian Area,</em> Data in Brief, 46 (2023), 108843, <a href="http://doi.org/10.1016/j.dib.2022.108843">10.1016/j.dib.2022.108843</a>, (<a href="https://www.sciencedirect.com/science/article/pii/S2352340922010460">https://www.sciencedirect.com/science/article/pii/S2352340922010460</a>)</p> <ol> </ol> <p> </p> <p><strong>Abstract</strong></p> <p>The <strong>MADIA gridded dataset</strong> provides the series of the main <strong>agro-meteorological </strong>variables derived from ERA5 hourly surface data, across the Italian domain for the period <strong>1981-2022</strong>, and their respective 1981-2010 and 1991-2020 <strong>climate normals</strong>,<strong> </strong>as well as the following statistics on the 30-year dekadal values of each variable: absolute minimum and maximum, 5<sup>th</sup>, 10<sup>th</sup>, 50<sup>th</sup>, 90<sup>th</sup>, 95<sup>th</sup> percentiles. Temporal and spatial resolutions are <strong>10-daily</strong> and <strong>0.25 degrees</strong> respectively. The dataset contains time series of minimum, average and maximum air temperature, minimum and maximum air relative humidity, wind speed, solar radiation, precipitation and reference evapotranspiration according to the FAO Penman-Monteith method. The dataset is provided in both <strong>NetCDF </strong>and <strong>csv </strong>format. In addition, discovery and description metadata are provided. In order to facilitate the data reuse for computing statistics at Italian <strong>NUTS 2 and 3</strong> levels, a complementary vector file is provided which reports the cell weight in terms of fraction covered of each administrative unit considered. Another vector file is included with the <strong>ERA5 cell polygons</strong> covering the Italian country for visualizing and mapping csv data. </p> <p>A <strong>daily version of the MADIA dataset</strong> (only in csv format) is also available on Zenodo at <a href="http://doi.org/10.5281/zenodo.7621453">https://doi.org/10.5281/zenodo.7621453</a>.</p> <p>Both MADIA datasets will be periodically updated.</p> <p><strong>Attached content</strong></p> <p>A ZIP archive composed by the following folders</p> <ol> <li>nc_data: annual time series from 1981 to 2022 and climate normals (1981-2010 and 1991-2020) in NetCDF format</li> <li>csv_data: annual time series from 1981 to 2022 and climate normals (1981-2010 and 1991-2020) in csv format</li> <li>metadata: discovery and description metadata </li> <li>shp_data: two complementary vector layers with the NUTS2-3 cover fractions and the ERA5 cell polygons for Italy</li> </ol> <p><strong>Acknowledgments</strong></p> <p>This work was supported by the Italian Ministry of Agricultural, Food and Forestry Policies (AgriDigit-Agromodelli, DM n. 36502 of 20/12/2018)</p>
ITTV - A Dataset of Italian Television for Automatic Genre Classification
<p>ITTV is a publicly available dataset of Italian TV programs introduced in </p> <blockquote> <p>Alessandro Ilic Mezza, Paolo Sani, and Augusto Sarti, "Automatic TV Genre Classification Based on Visually-Conditioned Deep Audio Features," in 2023 31st European Signal Processing Conference (EUSIPCO), 2023.</p> </blockquote> <p>ITTV consists of 2625 manually annotated YouTube videos, totaling over 670 hours. Each clip is assigned one of seven classes:</p> <ul> <li>Cartoons</li> <li>Commercials</li> <li>Football</li> <li>Music</li> <li>News</li> <li>Talk Shows</li> <li>Weather Forecast</li> </ul> <p>ITTV genre taxonomy is similar to that of the well-known RAI dataset described in</p> <blockquote> <p>Maurizio Montagnuolo and Alberto Messina, "Parallel neural networks for multimodal video genre classification,” Multimedia Tools and Applications, vol. 41, no. 1, pp. 125–159, 2009.</p> </blockquote> <p>The dataset contains genre annotations and metadata in CSV format. Please note that audio data is not provided.</p> <p>We provide the annotations for a balanced training (1575 clips) and validation (525 clips) split, as well as for a disjoint test set containing 525 installments from TV programs not included in the development set.</p> <p>As YouTube continuously updates, some videos may not be available in the future. Although we intend to keep ITTV updated as best as possible, please note that some content may not be available at any given time.</p> <p>Some YouTube videos (especially from the <code>Football</code> class and, to a lesser extent, the <code>Cartoons</code> class) may only be available in some countries due to regional restrictions imposed by the content creator. All videos are known to be accessible from Italy (last accessed on Nov. 25th, 2022.)<br> <br> Please contact Alessandro Ilic Mezza for further questions (e-mail: alessandroilic.mezza@polimi.it).</p>
Italian Polarity Lexicon
<p><strong>Italian polarity lexicon</strong></p> <p>Annotator: Susanna Tron</p> <p><br> Prefixes from the appraisal theory:</p> <p>A appreciation <br> F emotion<br> J judgement</p> <p><br> Polarity Labels:</p> <p>A_POS, F_POS, J_POS positive<br> A_NEG, F_NEG, J_NEG negative</p> <p>INT intensifier<br> DIM diminisher<br> SHI shifter</p> <p><br> Statistics:</p> <p>1'626 nouns (370 POS, 1242 NEG, 7 INT, 2 DIM, 5 SHI) <br> 1'549 adjectives (421 POS, 1060 NEG, 55 INT,8 DIM,5 SHI)<br> 206 adverbs (76 POS, 78 NEG, 44 INT, 5 DIM, 3 SHI)<br> 156 verbs (29 POS, 126 NEG, 0 INT, 1 DIM, 0 SHI)</p> <p>alltogether 3538 entries</p> <p><br> Format:</p> <p>lemma,polarity,part of speech</p> <p><br> Part of Speech:</p> <p>noun<br> adjective<br> adverb<br> verb</p> <p>Examples:</p> <p>allegrissimo,F_POS,adjective<br> allergia,A_NEG,noun<br> altamente,INT,adjective<br> alterigia,J_NEG,noun</p>
Italian Verb Lexicon for Sentiment Inference
<p><strong>Italian Verb Lexicon for Sentiment Inference</strong></p> <p><strong>Theory:</strong></p> <p>For a description of the theory behind the specifications of the corpus, please read the attached paper. </p> <p><br> <strong>Example of json entry:</strong></p> <p>{"verb": "soddisfare", "frames": [{"fillers": ["Subj", "DirObj/IndObj"], "polarity": "POS", "effects": [["DirObj/IndObj", "pos"]], "expectations": [], "examples": ["L'offerta ha soddisfatto i clienti.", "Soddisfare al pubblico."], "remarks": [], "relations": [["Subj", "DirObj/IndObj", "pro"]]}]}</p> <p><strong>Description: </strong><br> The verb "soddisfare" has 2 frames, a subject followed by a direct or indirect object. The verb polarity is positive. There is an positive effect on the direct (indirect) object. No expectations. There is a in favour (pro) relation from the subject to the direct (indirect) object. Two example sentenes are given.</p> <p><strong>Synonyms:</strong><br> Some entries are references to synonym verbs with identical frames:</p> <p>{"verb": "consacrare", "germanTranslation": "widmen", "frameReference": "dedicare", "examples": ["Consacrare tempo alle sue passioni"]}</p> <p>Here, "cosacrare" and "dedicare" are assumed synonyms with the same syntactic frames.</p> <p><strong>Used Tags:</strong></p> <p>A few explanations on the tags used in the verb specifications sheets.</p> <p>Subj = subject</p> <p>DirObj = direct object, as in "Il professore legge __il giornale__".</p> <p>IndObj = indirect object, as in "Permettere qualcosa __a qualcuno__".</p> <p>RefObj = reflexive object (pronoun), as in "La squadra avversaria __si__ è arrabbiata moltissimo". </p> <p>PrepObj[prep] = prepositional phrase; the preposition is specified in the square brackets. If more than one preposition can occur,<br> no specification is given.</p> <p>SubCl = a subordinate clause, usually introduced by "che" or "di" such as in "Ha detto __di andarsene__", <br> "Ha detto __che tutto è andato bene__".</p> <p>mod = any type of modifier, mostly adverbs, e.g. "Se ne è andato __subito__".</p> <p><br> *, e.g. mod* = indicates optionality</p> <p><br> </p> <p> </p>
Italian TikTok users online behaviour patterns and social attitudes (survey)
<p>Survey of 500 young TikTok users (18-35) in Italy covering online behaviour patterns and social attitudes</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.