Aller au contenu principal
Accès ouvert déclaré 2026 dataset

Unveiling Shortfalls in Biodiversity Knowledge of Brazilian Butterflies [Dataset]

0Citations signalées — pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Résumé fourni par la source

There are 5 files associated with this manuscript:(1) taxon_data.Rdata: This dataset (in R format) contains 10 columns with species-level information:"speciesName": currently valid species name based on the Taxonomic Catalogue of the Fauna of Brazil."authorship": species authorship."year": year of description."subfamily": taxonomic subfamily each species belong to."family": taxonomic family each species belong to."genus" taxonomic genus each species belong to."synonyms": unique synonyms associated with each valid species."subspecies": subspecies associated with each valid species."genbank_nseq": Number of genetic sequences available in the GenBank database for each species."bold_nseq": Number of genetic sequences available in the BOLD database for each species. (2) KnowledgeGapsBrButterflies.r: This R code is used to perform all analyses of our work as well as for creating the images on the main text and supplementary material. The script is commented by the authors and intuitively structured, including needed packages and their respective versions. (3) run_models.r: This R code is used to run the models used to estimate the Linnean shortfall for Brazilian butterflies. (4) functions.r: This R code contains user-created functions used in the main manuscript. (5) lepi_specieslink.txt: This file contains occurrence data for Brazilian butterflies downloaded from the speciesLink database. Code/Software: All analyses were performed using the software R version 4.5.1. Packages used and their respective versions can be found in the main R code used in this study. Usage notes: Software R is required to open the R codes and the associated Rdata files containing the datasets. The TXT file are the most basic text file format and are compatible with all plain text editors. Materials and MethodsTaxonomic DataWe used as our focal taxa 3,567 butterfly species listed in the Taxonomic Catalogue of the Fauna of Brazil (TCFB) ver. 1.23 [51], which also includes non-native species occurring in the country, such as Hypolimnas misippus from the Paleotropics, which naturally colonized the Caribbean region and now has established populations in northern Brazil (https://www.nic.funet.fi/index/Tree_of_life/insecta/lepidoptera/ditrysia/papilionoidea/nymphalidae/nymphalinae/hypolimnas/). The dynamic nature of taxonomic studies leads to changes in species names over time [e.g., 52,53]. To account for such changes when computing some proxies for the three shortfalls analysed here (see below), we compiled unique names and subspecies associated with each butterfly species, using data from the TCFB and the Catalogue of Life (CoL) Checklist ver. 2024.11. Unique names correspond to a single valid species, whereas ambiguous names are linked to multiple valid species.To describe temporal patterns in species descriptions, we fitted several models (linear, generalized additive model [GAM], second-order [quadratic] and third-order [cubic] polynomial regressions, and piecewise regression) to the annual number of species described as a function of year of description. We selected the best‐fitting model using the Akaike Information Criterion corrected for small sample sizes [AICc; 54] and used it to visualize and describe the data. In addition, we calculated mean annual species description rates for the entire study period and for each decade. Linnean ShortfallTo comprehend the magnitude of the Linnean shortfall, we estimated the number of undescribed butterfly species in Brazil using non-linear models based on species discovery curves derived from description dates (obtained from authority information available in the TCFB). These models consider taxonomic effort by incorporating the number of taxonomists involved in species description per time interval. The basic negative exponential model (ΔSt = k (Stot − St)) assumes three distinct parameters: ΔSt is the number of species described per time interval, Stot is the total number of species (described + undescribed), St is the cumulative number of species descriptions, and k represents a description efficiency parameter. When k is given as a function of time (k = Tt (a + bt)), where a and b are regression coefficients and T is the number of taxonomists who de-scribed species of the group per time interval, we have Joppa et al.’ model [55]. However, if k is given as a function of the number of species described per time interval (k = Tt (a + bΔSt)), we have Lu and He’s model [56]. Additionally, we fitted the logistic model using the gnls function from the nlme R package [57]. The residual variance was fitted as a power function of residual mean to account for over- or underdispersion. Models were ranked by Akaike Information Criterion (AICc), with the relative adequacy of each one evaluated by AICc weights (wAICc). Our analyses were conducted with data lumped in 5-year time intervals as proposed in the original code of Lu and He (2017) [56].It is important to emphasize that estimates of total species richness are dependent on the methods used for inference, which can produce highly variable results [see 58]. In particular, estimates based on temporal trends in species descriptions are often associated with substantial uncertainty, especially for poorly studied taxonomic groups [59]. Some authors even argue that the many uncertainties surrounding the species discovery process make the extrapolation of species accumulation curves inherently unreliable for estimating total species richness [60,61]. Nevertheless, establishing a baseline estimate re-mains valuable for communicating biodiversity challenges to the general public and policymakers, as well as for illustrating the scale of the taxonomic work still ahead. Regardless of these uncertainties, one conclusion remains clear: if we aim to uncover Earth’s hidden diversity—and its potential ecological, evolutionary, and societal value—before it is lost through ignorance, we will still need substantially greater investment in taxonomic research and field exploration, or, in Edward O. Wilson’s words, “more boots on the ground” [62]. Wallacean ShortfallTo estimate the lack of information on Brazilian butterflies’ geographic distribution, we downloaded on December 09, 2025 all available lepidopteran occurrence records from Brazil on both GBIF (98,015 records; https://doi.org/10.15468/dl.jdu6xr) and speciesLink (69,607 records) databases. We then retained only those belonging to butterfly species, with geographic coordinates and no suspicious flags (i.e., potentially erroneous records) and removed data for human observations (e.g. iNaturalist data) due to high taxonomic uncertainties in species identifications [e.g., 63,64]. Additionally, we applied data cleaning procedures, such as removing duplicates (i.e., same specimen and locality for a given species) and records on country centroids using the CoordinateCleaner R package [65], yielding a total of 53,638 cleaned records in GBIF and 27,894 in speciesLink. We then combined this data and, after removing duplicates and records without species-level identification (e.g. Cissia sp.), we had a total of 74,464 records of Brazilian butterflies. We used this data to obtain summary statistics per taxa as well as sampling completeness (coverage) metrics [66] for each spatial unit, defined as grid cells of 1° resolution (≈ 110 × 110 km).To estimate the spatial distribution of butterfly species richness in Brazil, we used geostatistical interpolation methods [67–69], which take advantage of the spatial auto-correlation in geographical data, including species richness and its association with environmental variables [70]. Specifically, we used a regression-kriging (R-K) model to create continuous maps of species richness, which requires a subsample of data that are spatially autocorrelated and covary with other factors (e.g., environmental variables) throughout a target region [71]. For constructing this model, we first created a grid system over Brazil using a

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.