Summarizing Web Archive Corpora Via Social Media Storytelling By Automatically Selecting and Visualizing Exemplars

Document Type

Article

Publication Date

2023

DOI

10.1145/3606030

Publication Title

ACM Transactions on the Web

Volume

Article in Press

Pages

1-48

Abstract

People often create themed collections to make sense of an ever-increasing number of archived web pages. Some of these collections contain hundreds of thousands of documents. Thousands of collections exist, many covering the same topic. Few collections include standardized metadata. This scale makes understanding a collection an expensive proposition. Our Dark and Stormy Archives (DSA) five-process model implements a novel summarization method to help users understand a collection by combining web archives and social media storytelling. The five processes of the DSA model are: select exemplars, generate story metadata, generate document metadata, visualize the story, and distribute the story. Selecting exemplars produces a set of k documents from the N documents in the collection, where k < <N, thus reducing the number of documents visitors need to review to understand a collection. Generating story and document metadata selects images, titles, descriptions, and other content from these exemplars. Visualizing the story ties this metadata together in a format the visitor can consume. Without distributing the story, it is not shared for others to consume. We present a research study demonstrating that our algorithmic primitives can be combined to select relevant exemplars that are otherwise undiscoverable using a conventional search engine and query generation methods. Having demonstrated improved methods for selecting exemplars, we visualize the story. Previous work established that the social card is the best format for visitors to consume surrogates. The social card combines metadata fields, including the document’s title, a brief description, and a striking image. Social cards are commonly found on social media platforms. We discovered that these platforms perform poorly for mementos and rely on web page authors to supply the necessary values for these metadata fields. With web archives, we often encounter archived web pages that predate the existence of this metadata. To generate this missing metadata and ensure that storytelling is available for these documents, we apply machine learning to generate the images needed for social cards with a Precision@1 of 0.8314. We also provide the length values needed for executing automatic summarization algorithms to generate document descriptions. Applying these concepts helps us create the visualizations needed to fulfill the final processes of story generation. We close this work with examples and applications of this technology.

Rights

© 2023 Copyright held by the owner/author(s). Publication rights licensed to ACM.

ACM acknowledges that this contribution was authored or co-authored by an employee, contractor or affiliate of the United States government. As such, the Government retains a nonexclusive, royalty-free right to publish or reproduce this article, or to allow others to do so, for Government purposes only.

"ACM treats links as citations (references to objects) rather than as incorporations (embedding of objects). Permission is not needed to create links to citations in The ACM Digital Library or Online Guide to Computing Literature. ACM encourages the widespread distribution of links to the definitive Version of Records of its copyrighted works in the ACM Digital Library and does not require that authors obtain prior permission to include such links in their new works.

However, someone who creates a work or a service whose pattern of links substantially duplicates an ACM-copyrighted volume or issue should get prior permission from ACM. One example: the creator of "A Table of Contents for the Current Issue of TODS" -- consisting of citations and active links to author-versions of the works in the latest issue of TODS -- needs ACM permission because that creator is reproducing an ACM-copyrighted work. If all the links in the "Table of Contents" pointed to the ACM-held definitive Version of Records, ACM would normally give permission because then the new work advertises an ACM work. To avoid misunderstandings, consult with ACM before duplicating an ACM work via links.

If an author wishes to embed a copyrighted object---rather than a link---in a new work, that author needs to obtain the copyright holder's permission."

Original Publication Citation

Jones, S. M., Klein, M., Weigle, M. C., & Nelson, M. L. (2023). Summarizing web archive corpora via social media storytelling by automatically selecting and visualizing exemplars. ACM Transactions on the Web. Advance online publication. https://doi.org/10.1145/3606030

ORCID

0000-0002-4372-870X (Jones), 0000-0003-0130-2097 (Klein), 0000-0002-2787-7166 (Weigle), 0000-0003-3749-8116 (Nelson)

Share

COinS