Geology ReportsSearch

USGS · 70277088

The digital archivist: Automating legacy macroseismic data processing using large language models

Abstract

Macroseismic data are a key resource to investigate shaking and damage from preinstrumental and early instrumental eras. However, data are often stored as inconsistently formatted reports describing observed shaking and damage, making manually parsing and interpreting accounts labor‐intensive. We introduce a novel workflow using Google’s Gemini 2.5 Pro large language model (LLM) to automate the extraction and structuring of macroseismic observations from summary reports. We apply this workflow to the 22 March 1957 M 5.3 Daly City, California, earthquake as a case study. We used Gemini to extract addresses, originally assigned modified Mercalli intensity values, and descriptions from each report. To address coordinate precision limits, addresses were geocoded via Google’s Geocoding application programming interface. This workflow yielded over 2300 geocoded intensity reports for the Daly City earthquake. We use the geocoded accounts, with the original report intensity assignments, to develop a shaking intensity map that in some respects rivals modern Did You Feel It? Maps. We also extract and present data for the 9 February 1971 M L 6.7 Sylmar, California, earthquake. Our results demonstrate the potential of LLMs for reliably extracting and analyzing large, unstructured macroseismic datasets. LLMs offer a scalable solution for rapidly digitizing macroseismic archives, enabling their broader use to constrain ground‐motion models in modern seismic hazard analysis and to improve our understanding of site effects in urban areas. The concepts explored here may also be applied to the handling of other legacy seismological and earth science data.

Explore related subjects

Keep this discovery

BibTeXRIS

Aarnav Agrawal, Susan E. Hough, Mostafa Mousavi, Margaret Hellweg, William Ellsworth, Clara Yoon, Salvador Blanco. 2026-07-01. The digital archivist: Automating legacy macroseismic data processing using large language models. https://doi.org/10.1785/0220250362

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

Digitizer Suite: The Albuquerque Seismological Laboratory Digitizer Testing Suite

Laboratory testing of digitizers and seismometers helps ensure that prior to deployment the instrumentation can produce high quality data and is operating within specifications. In this work we detail the software package called: the Albuquerque Seismological Laboratory (ASL) Digitizer Test Suite. This Java software package provides several algorithms to verify various performance parameters of digitizers commonly used for recording analog seismic instruments. The goal of these tests is not to be exhaustive, but to identify common failures that could compromise the integrity of seismic data being recorded on the digitizer. For example, Sandia National Laboratories (e.g., Slad and Merchant, 2018) routinely do comprehensive testing of digitizers for various monitoring missions. While these tests reports are valuable for comprehensively characterizing a recording system, it would be resource intensive to conduct such tests on every seismic recorder used in a network. We focus on tests that include ways to estimate the sensitivity, timing, self-noise, and clip-level of the digitizer, as well as the fidelity of the signal being recorded. The software is publicly available and provides a way for the community to verify the integrity of a digitizer using a minimum amount of outside equipment.

Seismological Research Letters

pySATSI: A Python package for computing focal mechanism stress inversions

We introduce pySATSI, a Python package for computing earthquake focal mechanism stress inversions. This algorithm can handle a wide variety of types of stress inversion problems with a single script and can duplicate many capabilities of preceding methodologies. We also add new capabilities that include spatiotemporally variable inversion grids, damped stress estimates for clusters with few or no focal mechanisms, and variable fault‐plane ambiguities that the user can assign to individual events. In addition, we added the ability to use damped stress inversions with fault‐plane ambiguity probabilities that are weighted by fault instabilities. Our algorithm is computationally efficient with faster runtimes than previous algorithms, scales well for large datasets, and can be easily parallelized.

Seismological Research Letters