To what lengths do the CDS & NED go to harvest useful data from astronomical papers? Marion Schmitz –
[email protected] on behalf of the NED team
[email protected] Emmanuelle Perret –
[email protected] on behalf of the CDS team
[email protected]
CDS & NED similarities Selected relevant data: Identifiers, Coordinates, Velocities, Object types, Sizes, References, Photometry, Wavelength & Accuracy of measurements... Tools to help: ML & MatchEx... (NED) DJIN / COSIM... (SIMBAD) anafile; setUCD… (VizieR) Automated verifications & Verifications from expertise At diverse steps in the process a Bibcodes © LISA VIII
To what lengths does the CDS go to harvest useful data from astronomical papers? LISA X – 2025 November, 5th I. Selection of data: ➢Evolution of selection criteria ➢Quantity vs Quality II. The other side of the scene: ➢Value-added data in VizieR ➢Will AI change the way we work?
I. Selection of data – How do we process astronomers’ data at ? ●Read through 19 astronomical journals (9 with priority 1 including ApJ, ApJS, AJ, A&A, MNRAS, Nature…) ●Standardize large tables, describe data: ReadMe file & metadata, and verify data ●Cross-identifications between astronomical objects from tables and SIMBAD Status at CDS
I. Selection of data – historical evolution for (SIMBAD) ●Before 2013, 63 astronomical journals processed (28 in priority 1 or 2) ●Now, 19 journals in priority 1 or 2. ●Go through ALL articles – update astronomical objects in text, figures, small tables, etc. (Aline’s talk) ●Tag(s): - large tables (>100 rows) for VizieR, - other tables (too large for DJIN) for SIMBAD’s 2d team (COSIM), - new astronomical objects for acronym’s creation (Dictionary) Publications & archives COSIM VizieR 1st SIMBAD team (DJIN) Stat facts 4 documentalists for ~14,000 ref./yr
I. Selection of data – historical evolution for COIM Stat facts ~10% of references include tables ~1200 catalogs are ingested, every year, by 3 documentalists Publications & archives VizieR 1st SIMBAD team (DJIN) ●Selection of papers: ➢Before 2019, all tables >100 rows were processed… ➢Now, selection criteria: - no more atomic data, simulations... - Priority for new observations, etc. ●Selection of tables in a paper: ➢Following to the goal of the paper or ➢Allowing links toward other data (e.g.: observation logs for spectra…) ➢Allowing to update SIMBAD. Focus on tables (selected by the DJIN team – mainly)
I. Selection of data – historical evolution for (SIMBAD) COSIM ●Before 2006, every table with astronomical objects coming from VizieR. ●In 2006, one weekly meeting with astronomers (at least 2) and a documentalist to determine priorities and which data, with which qualities taken into account. ●Since 2013, bi-weekly meeting. Scientific criteria determined by CDS’ astronomers (specialized in diverse fields) Publications & archives COSIM VizieR 1st SIMBAD team (DJIN) Stat facts ●Depending of the year of publication, between 13% up to ~58% (max. in 2007) of tables are degraded to a secondary priority. ●On average, since 2013, ~490 tables (from VizieR or the publication directly) are processed every year by 3 documentalists. Focus on tables
I. Selection of data – Quantity vs Quality, the hard balance COSIM from DJIN Publications (including archives like Zenodo...) COSIM Time to process data Number of references QUANTITY QUALITY
To what lengths does the CDS go to harvest useful data from astronomical papers? LISA X – 2025 November, 5th I. Selection of data: ➢Evolution of selection criteria ➢Quantity vs Quality II. The other side of the scene: ➢Value-added data in VizieR ➢Will AI change the way we work?
UCD Finder cds.unistra.fr/UCD/ ucd-finder/ II. The other side of the scene – How the ReadMefile is also VizieR – Tool example 2: setUCD: internal tool = suggestions of Unified Content Descriptors VizieR ReadMe to...
II. The other side of the scene – The hidden file «.status»: to customize data & add more metadata/data in VizieR Also VizieR... Column «Field» added by CDS Column «FileName» added by CDS
II. The other side of the scene – To what lengths is subjective & depends on time Fun fact A new documentalist made a Python script to plot a histogram...
II. The other side of the scene – Sometimes users ask us to go far... Spectra added in 2025... Fun fact In 1999, the .status file to ingest tables in VizieR contained ~40 rows; now it is ~150 rows.
II. The other side of the scene – Will AI change the way we work? ●The first question to ask about AI concerns ethical issues. We cannot, as a research institute, use it without questioning this. This question, in my opinion, should come first. ●The second question is what are the goals of using AI. Is it a gain of time to add more and more data in the databases (quantity) or do we want to improve quality? ●That being said, there are several places where we could use AI for VizieR and we did some tests but thinking more about it is necessary (pending project): ➢Save time on descriptions when it is just a matter of formatting data. One test with Claude => too soon to tell... ➢It has been used to convert a table from an image into an ASCII table (but this case is rare now)… - ChatGPT advice => Python library: pytesseract ➢Maybe it will help to keep track of facilities and help us with the associated data (tests in progress with DeepSeek, ChatGPT, Gemini and Grok) ➢It could be nice to replace the setUCD program…
In a nutshell... Description / Indexation / Verification: ➢Short Title ➢Keywords for tables ➢Observations/Origin of samples ➢Standard formats & Units & Labels Columns, tables or files added by CDS Content of each column of each table is verified .status file Display / Indexation / Verification: ➢Formats of values; default columns... ➢Metadata (UCDs, Filters, Dates...) ➢Additional columns & plots (Positions, SIMBAD, NED, LCs, spectra...) ➢Links between tables or toward cited catalogs ➢Corrections Selection Research Scientists Data curators Software engineers F T P Table 1 ReadMe file Standards Table 2
1 To what lengths does NED go to harvest useful data from astronomical papers? LISA X – 2025 November, 5th I. Selection of data: ➢ Journals covered ➢ Relevant data II. Triage: ➢ ML classification of articles (NAP) ➢ Names-only articles ➢ Articles with relevant data ➢ Very Large tables/catalogs (VLC) III. Data extraction: ➢ Text – copy ‘n paste/{ai4ai} ➢ Tables – VizieR/MRT/CSV/copy ‘n paste/{ai4ai} https://ned.ipac.caltech.edu IV. Data validation - I: ➢ NED interface ➢ Consistency tools - python V. Data validation - II: ➢ MatchEx ➢ Consistency checks - human VI. DB Update: ➢ PostgreSQL VII. Public Release: ➢ Data release ➢ Software release
2 NED: I. Selection of Data LISA X – 2025 November, 5thhttps://ned.ipac.caltech.edu ➢ Journals covered: •ApJ, ApJL, ApJS, AJ, A&A, MNRAS, Nature ➢ Relevant data •Names •Object type – Galaxy, QSO, Cluster of Galaxies, … •Position + uncertainty •Redshift + uncertainty
3 NED: II. Triage LISA X – 2025 November, 5thhttps://ned.ipac.caltech.edu ➢ ML classification of articles •Chen et al. 2022PASP..134a4501C •NED appropriate papers (NAP) determined by the ML classifier (Stanford) •NAP flagged for positions, redshifts, photometry, NED acknowledgments ➢ Names-only articles •Names are normalized to NED style via NED’s name resolver (e.g. M31 MESSIER 031) •Loaded into DB without further investigation ➢ Articles with relevant data •Analyzed and prepared by Production Team member •Full validation and loading procedures by same Prod team member ➢ Very Large tables/catalogs (VLC) •Tables with >10,000 entries put in BackBurner list or as resources permit •Processed in total automation mode after spot-checking of test areas by human
4 NED: III. Data Extraction LISA X – 2025 November, 5thhttps://ned.ipac.caltech.edu ➢ Text •copy ‘n paste from PDF or HTML •{ai4ai} Artificial Intelligence for Archive Ingest –in development ➢ Tables •VizieR (if available) •AAS MRT (Machine Readable Table) •CSV (GitHub, Zenodo, author provided, …) •copy ‘n paste •{ai4ai} – in development