scieee AI-readable full text Open interactive document viewer

To what length the CDS goes to harverst useful data from astronomical papers?

Perret, Emmanuelle; Schmitz, Marion

Abstract

This interesting question was submitted to us very recently by an astronomer discovering the tables in his paper online in VizieR, one of the CDS (Strasbourg astronomical Data Center) databases, even though he had not submitted them himself. We also receive many questions from authors on the hotline asking why their data are not included in the CDS databases. The purpose of this presentation is to look at the other side of the scene: putting tables online in VizieR is not a simple, magical process where you press a button and everything is displayed and made available according to the FAIR (Findable, Accessible, Interoparable and Reusable) principles. This is especially true given the goal of providing high-quality data –e.g. VizieR is certified by the CoreTrustSeal-- and the pressure of constantly increasing data volumes. Another recurring interesting (and provocative) question can be addressed: why are the documentalists –who feed the CDS databases on a daily basis-- not yet being replaced by artificial intelligence? It is very early to answer this question, but we can start by looking at what is currently being done through real-life examples and see if improvments are possible. This presentation will focus mainly on VizieR (the catalog database), but at the CDS, everything is interconnected, whether it be the various services provided or the different professions within the team providing these services.

Full text

To what lengths do the CDS & NED go to harvest useful data from astronomical papers? Marion Schmitz – [email protected] on behalf of the NED team [email protected] Emmanuelle Perret – [email protected] on behalf of the CDS team [email protected] CDS & NED similarities Selected relevant data: Identifiers, Coordinates, Velocities, Object types, Sizes, References, Photometry, Wavelength & Accuracy of measurements... Tools to help: ML & MatchEx... (NED) DJIN / COSIM... (SIMBAD) anafile; setUCD… (VizieR) Automated verifications & Verifications from expertise At diverse steps in the process a Bibcodes © LISA VIII To what lengths does the CDS go to harvest useful data from astronomical papers? LISA X – 2025 November, 5th I. Selection of data: ➢Evolution of selection criteria ➢Quantity vs Quality II. The other side of the scene: ➢Value-added data in VizieR ➢Will AI change the way we work? I. Selection of data – How do we process astronomers’ data at ? ●Read through 19 astronomical journals (9 with priority 1 including ApJ, ApJS, AJ, A&A, MNRAS, Nature…) ●Standardize large tables, describe data: ReadMe file & metadata, and verify data ●Cross-identifications between astronomical objects from tables and SIMBAD Status at CDS I. Selection of data – historical evolution for (SIMBAD) ●Before 2013, 63 astronomical journals processed (28 in priority 1 or 2) ●Now, 19 journals in priority 1 or 2. ●Go through ALL articles – update astronomical objects in text, figures, small tables, etc. (Aline’s talk) ●Tag(s): - large tables (>100 rows) for VizieR, - other tables (too large for DJIN) for SIMBAD’s 2d team (COSIM), - new astronomical objects for acronym’s creation (Dictionary) Publications & archives COSIM VizieR 1st SIMBAD team (DJIN) Stat facts 4 documentalists for ~14,000 ref./yr I. Selection of data – historical evolution for COIM Stat facts ~10% of references include tables ~1200 catalogs are ingested, every year, by 3 documentalists Publications & archives VizieR 1st SIMBAD team (DJIN) ●Selection of papers: ➢Before 2019, all tables >100 rows were processed… ➢Now, selection criteria: - no more atomic data, simulations... - Priority for new observations, etc. ●Selection of tables in a paper: ➢Following to the goal of the paper or ➢Allowing links toward other data (e.g.: observation logs for spectra…) ➢Allowing to update SIMBAD. Focus on tables (selected by the DJIN team – mainly) I. Selection of data – historical evolution for (SIMBAD) COSIM ●Before 2006, every table with astronomical objects coming from VizieR. ●In 2006, one weekly meeting with astronomers (at least 2) and a documentalist to determine priorities and which data, with which qualities taken into account. ●Since 2013, bi-weekly meeting. Scientific criteria determined by CDS’ astronomers (specialized in diverse fields) Publications & archives COSIM VizieR 1st SIMBAD team (DJIN) Stat facts ●Depending of the year of publication, between 13% up to ~58% (max. in 2007) of tables are degraded to a secondary priority. ●On average, since 2013, ~490 tables (from VizieR or the publication directly) are processed every year by 3 documentalists. Focus on tables I. Selection of data – Quantity vs Quality, the hard balance COSIM from DJIN Publications (including archives like Zenodo...) COSIM Time to process data Number of references QUANTITY QUALITY To what lengths does the CDS go to harvest useful data from astronomical papers? LISA X – 2025 November, 5th I. Selection of data: ➢Evolution of selection criteria ➢Quantity vs Quality II. The other side of the scene: ➢Value-added data in VizieR ➢Will AI change the way we work? UCD Finder cds.unistra.fr/UCD/ ucd-finder/ II. The other side of the scene – How the ReadMefile is also VizieR – Tool example 2: setUCD: internal tool = suggestions of Unified Content Descriptors VizieR ReadMe to... II. The other side of the scene – The hidden file «.status»: to customize data & add more metadata/data in VizieR Also VizieR... Column «Field» added by CDS Column «FileName» added by CDS II. The other side of the scene – To what lengths is subjective & depends on time Fun fact A new documentalist made a Python script to plot a histogram... II. The other side of the scene – Sometimes users ask us to go far... Spectra added in 2025... Fun fact In 1999, the .status file to ingest tables in VizieR contained ~40 rows; now it is ~150 rows. II. The other side of the scene – Will AI change the way we work? ●The first question to ask about AI concerns ethical issues. We cannot, as a research institute, use it without questioning this. This question, in my opinion, should come first. ●The second question is what are the goals of using AI. Is it a gain of time to add more and more data in the databases (quantity) or do we want to improve quality? ●That being said, there are several places where we could use AI for VizieR and we did some tests but thinking more about it is necessary (pending project): ➢Save time on descriptions when it is just a matter of formatting data. One test with Claude => too soon to tell... ➢It has been used to convert a table from an image into an ASCII table (but this case is rare now)… - ChatGPT advice => Python library: pytesseract ➢Maybe it will help to keep track of facilities and help us with the associated data (tests in progress with DeepSeek, ChatGPT, Gemini and Grok) ➢It could be nice to replace the setUCD program… In a nutshell... Description / Indexation / Verification: ➢Short Title ➢Keywords for tables ➢Observations/Origin of samples ➢Standard formats & Units & Labels Columns, tables or files added by CDS Content of each column of each table is verified .status file Display / Indexation / Verification: ➢Formats of values; default columns... ➢Metadata (UCDs, Filters, Dates...) ➢Additional columns & plots (Positions, SIMBAD, NED, LCs, spectra...) ➢Links between tables or toward cited catalogs ➢Corrections Selection Research Scientists Data curators Software engineers F T P Table 1 ReadMe file Standards Table 2 1 To what lengths does NED go to harvest useful data from astronomical papers? LISA X – 2025 November, 5th I. Selection of data: ➢ Journals covered ➢ Relevant data II. Triage: ➢ ML classification of articles (NAP) ➢ Names-only articles ➢ Articles with relevant data ➢ Very Large tables/catalogs (VLC) III. Data extraction: ➢ Text – copy ‘n paste/{ai4ai} ➢ Tables – VizieR/MRT/CSV/copy ‘n paste/{ai4ai} https://ned.ipac.caltech.edu IV. Data validation - I: ➢ NED interface ➢ Consistency tools - python V. Data validation - II: ➢ MatchEx ➢ Consistency checks - human VI. DB Update: ➢ PostgreSQL VII. Public Release: ➢ Data release ➢ Software release 2 NED: I. Selection of Data LISA X – 2025 November, 5thhttps://ned.ipac.caltech.edu ➢ Journals covered: •ApJ, ApJL, ApJS, AJ, A&A, MNRAS, Nature ➢ Relevant data •Names •Object type – Galaxy, QSO, Cluster of Galaxies, … •Position + uncertainty •Redshift + uncertainty 3 NED: II. Triage LISA X – 2025 November, 5thhttps://ned.ipac.caltech.edu ➢ ML classification of articles •Chen et al. 2022PASP..134a4501C •NED appropriate papers (NAP) determined by the ML classifier (Stanford) •NAP flagged for positions, redshifts, photometry, NED acknowledgments ➢ Names-only articles •Names are normalized to NED style via NED’s name resolver (e.g. M31 MESSIER 031) •Loaded into DB without further investigation ➢ Articles with relevant data •Analyzed and prepared by Production Team member •Full validation and loading procedures by same Prod team member ➢ Very Large tables/catalogs (VLC) •Tables with >10,000 entries put in BackBurner list or as resources permit •Processed in total automation mode after spot-checking of test areas by human 4 NED: III. Data Extraction LISA X – 2025 November, 5thhttps://ned.ipac.caltech.edu ➢ Text •copy ‘n paste from PDF or HTML •{ai4ai} Artificial Intelligence for Archive Ingest –in development ➢ Tables •VizieR (if available) •AAS MRT (Machine Readable Table) •CSV (GitHub, Zenodo, author provided, …) •copy ‘n paste •{ai4ai} – in development