Preservation and Trust in Scholarly Communications, Part One: Preservation, Access, and Trust
Scope
How do we keep collections safe, accessible, and credible when funding, policy, and public sentiment pull in different directions? We’ll explore governance, provenance, versioning/audit trails, community input, and transparency practices that strengthen trust while resisting censorship.
Confirmed Speakers: Micah Altman, Social and Information Scientist, MIT; Kate Murray, Digital Projects Coordinator, Library of Congress; and Kate Wittenberg, Managing Director, Portico. Trevor Owens, Chief Research Officer at AIP has advised the shaping of this program and will serve as the moderator.
Event Sessions
Speakers
Trevor Owens, Chief Research Officer for AIP Publishing served as the moderator for this program.
The following questions were posed to our speakers:
What are the most pressing concerns in digital preservation right now?
How is AI changing what we need to preserve—and what we need to know about what we preserve?
What aspects of trustworthiness are coming under strain, and how can we continue to trust enduring access to digital material?
How do institutions navigate dependence on increasingly centralized commercial infrastructure while maintaining resilient preservation?
How should preservation respond when AI scraping, rights, access, and openness begin pulling in different directions?
What are you still thinking or “ruminating” about as a result of this conversation?
Related Information and Shared Resources:
How Claude marks AI-generated content - Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, as a provider of both generative AI models and generative AI systems. This article describes how we’re planning to put those commitments into practice, how marking works, and what its limitations are.
Generative AI for Trustworthy, Open, and Equitable Scholarship -This paper provides a roadmap for leveraging Generative AI to address known challenges to research integrity, focusing on potential innovations and interventions in peer review, open data sharing, accessibility, and inclusion. By Chris Bourg, Sue Kriegsman, Nick Lindsay, Heather Sardis, Erin Stalberg, and Micah Altman
How Open Must Language Models be to Enable Reliable Scientific Inference? By James A. Michaelov, Catherine Arnett, Tyler A. Chang, Pamela D. Rivière, Samuel M. Taylor, Cameron R. Jones, Sean Trott, Roger P. Levy, Benjamin K. Bergen, Micah Altman - How does the extent to which a model is open or closed impact the scientific inferences that can be drawn from research that involves it? In this paper, we analyze how restrictions on information about model construction and deployment threaten reliable inference. We argue that current closed models are generally ill-suited for scientific purposes, with some notable exceptions, and discuss ways in which the issues they present to reliable inference can be resolved or mitigated. We recommend that when models are used in research, potential threats to inference should be systematically identified along with the steps taken to mitigate them, and that specific justifications for model selection should be provided.
Content Authenticity and Provenance in the Age of Artificial Intelligence - This white paper is intended for preservation administrators and practitioners across the LAMs sector who, in one way or another, find themselves addressing the impact AI technologies are having on the content integrity of their collections.”
Coalition for Content Authenticity and Provenance Content Credentials - The Coalition for Content Provenance and Authenticity (C2PA) addresses the prevalence of misleading information online through the development of technical standards for certifying the source and history (or provenance) of media content. C2PA is a Joint Development Foundation project.
Expertise’ shouldn’t be a bad word – expert consensus guides science and society - Expert consensus guides science and society.
Views from the front lines of Trump’s war on the science community: by Micah Altman and Philip N. Cohen, opinion contributors - The Trump administration has unleashed a tsunami of budget cuts to federal science programs. Mass firings have taken place at both the Department of Health and Human Services and the Department of Education, part of a deliberate decimation of research staff across the federal government.
The Scholarly Knowledge Ecosystem: Challenges and Opportunities for the Field of Information by Micah Altman and Philip N. Cohen - The scholarly knowledge ecosystem presents an outstanding exemplar of the challenges of understanding, improving, and governing information ecosystems at scale. This article draws upon significant reports on aspects of the ecosystem to characterize the most important research challenges and promising potential approaches. The focus of this review article is the fundamental scientific research challenges related to developing a better understanding of the scholarly knowledge ecosystem.
Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems - Providers shall ensure that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system, unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect, taking into account the circumstances and the context of use. This obligation shall not apply to AI systems authorised by law to detect, prevent, investigate or prosecute criminal offences, subject to appropriate safeguards for the rights and freedoms of third parties, unless those systems are available for the public to report a criminal offence.
SB-942 California AI Transparency Act (CAITA) - Existing law requires the Secretary of Government Operations to develop a coordinated plan to, among other things, investigate the feasibility of, and obstacles to, developing standards and technologies for state departments to determine digital content provenance. For the purpose of informing that coordinated plan, existing law requires the secretary to evaluate, among other things, the impact of the proliferation of deepfakes, defined to mean audio or visual content that has been generated or manipulated by artificial intelligence that would falsely appear to be authentic or truthful and that features depictions of people appearing to say or do things they did not say or do without their consent, on state government, California-based businesses, and residents of the state.
Making AI Use of Scholarly Content Traceable, Measurable, and Trustworthy: A Meeting Report from Cambridge Scholarly AI Workshop Making AI Use of Scholarly Content Traceable, Measurable, and Trustworthy: A Meeting Report from Cambridge Scholarly AI Workshop: By Todd A Carpenter, Tasha Mellins-Cohen, Monica Westin - This is the second in a series of posts on AI systems, provenance tracking in generative artificial intelligence systems, and the implications on usage and assessment. The first post was published last week. The next piece in this series will cover forthcoming community work related to provenance tracking and usage.
Tiered Community Recommendations for Content Authenticity and Provenance (TCR4CAP): Audio-Visual Working Group - Initiated in early 2026, the FADGI AV Working Group has established a new action team to develop tiered community recommendations for content authenticity and provenance (CAP) for digital audiovisual collections. The effort, inspired by the NDSA Levels of Digital Preservation, aims to help government, library, archive, and museum institutions determine practical, resource-appropriate approaches to documenting authenticity—especially in an era where AI increasingly interacts with institutional collections. The project is defining levels of practice ranging from basic integrity checks to more advanced implementations such as embedded provenance metadata or trust-center integrations. The framework is not intended to certify files or systems, but to support institutional planning, policy development, and shared community understanding.
A draft for public comment is expected to be released in summer 2026.
Selecting Efficient and Reliable Preservation Strategies: Modeling Long-term Information Integrity Using Large-scale Hierarchical Discrete Event Simulation by Micah Altman and Richard Landau - This article addresses the problem of formulating efficient and reliable operational preservation policies that ensure bit-level information integrity over long periods, and in the presence of a diverse range of real-world technical, legal, organizational, and economic threats. We develop a systematic, quantitative prediction framework that combines formal modelling, discrete-event-based simulation, hierarchical modelling, and then use empirically calibrated sensitivity analysis to identify effective strategies. Specifically, the framework formally defines an objective function for preservation that maps a set of preservation policies and a risk profile to a set of preservation costs, and an expected collection loss distribution.
Information wants someone else to pay for it: Laws of information economics and scholarly publishing by Micah Altman and Marguerite Aver - The increasing volume and complexity of research, scholarly publication, and research information puts an added strain on traditional methods of scholarly communication and evaluation. Information goods and networks are not standard market goods – and so we should not rely on markets alone to develop new forms of scholarly publishing. The affordances of digital information and networks create many opportunities to unbundle the functions of scholarly communication – the central challenge is to create a range of new forms of publication that effectively promote both market and collaborative ecosystems.
The NDSA Agenda for Digital Stewardship - The NDSA Agenda is a comprehensive overview of the state of global digital preservation. It casts its eye over current research trends, grants, projects, and various efforts spanning the preservation ecosystem. The agenda identifies successes and ongoing challenges in addition to providing some tangible recommendations to both researcher and practitioner alike. As both an overview and comprehensive dive into digital preservation issues, the audience ranges from high level to hands on experts. Funders can use this report as a signpost for the overall state of the profession.
Interventions in scholarly communication: Design lessons from public health by Micah Altman, Philip N. Cohen, and Jessica Polka - Many argue that swift and fundamental interventions in the system of scholarly communication are needed. However, there are substantial disagreements over the short- and long-term benefits of most proposed approaches to changing the practice of science communication, and the lack of systematic, empirically based research in this area makes these controversies difficult to resolve. We argue that experience within public health can be usefully applied to scholarly communication.
Guidance for Artificial Intelligence and Machine Learning - The C2PA specification can be used to add cryptographic information to detect the tampering of media files and streams. Similarly, the C2PA framework can also be used in artificial intelligence (AI) and machine learning (ML) systems to indicate the tampering of datasets, software, and models which are utilized during training and inference. This document provides guidance about how C2PA’s Content Credentials can be employed by AI and ML systems.
Creative Commons on the Signals project about alerting authors’ desires around AI training - AI is being built on the largest unregulated extraction of knowledge in history. The datasets powering today’s AI systems are drawn largely from publicly available information: research papers, educational resources, cultural works, images, and data shared by millions of people and institutions around the world. A significant portion of this knowledge exists because of CC licenses and the global movement for open sharing.
Additional Information
NISO assumes organizations register as a group. The model assumes that an unlimited number of staff will be watching the live broadcast in a single location, but also includes access to an archived recording of the event for those who may have timing conflicts.
Educational program contacts and registrants receive sign-on instructions via email three business days prior to the virtual event. If you have not received your instructions by the day before an event, please contact NISO headquarters for assistance via email (nisohq@niso.org).
Registrants for an event may cancel participation and receive a refund (less $30.00) if the notice of cancellation is received at NISO HQ (nisohq@niso.org) one full week prior to the event date. If received less than 7 days before, no refund will be provided.
Links to the archived recording of the broadcast are distributed to registrants 24-48 business hours following the close of the live event. Access to that recording is intended for internal use of fellow staff at the registrant’s organization or institution. Shared resources are posted to the NISO event page.
All events follow the NISO Code of Conduct. More information can be found here.
Broadcast Platform
NISO uses the Zoom platform for the purpose of broadcasting our live events. Zoom provides apps for a variety of computing devices (tablets, laptops, etc.) To view the broadcast, you will need a device that supports the Zoom app. Attendees may also choose to listen just to audio on their phones. Sign-on credentials include the necessary dial-in numbers, if that is your preference. Once notified of their availability, recordings may be viewed from the Zoom platform.
Event Dates
–
Fees
Designated educational program contacts at NISO member organizations automatically received sign-on credentials for regularly scheduled webinar events as a benefit of membership. If you are unsure who your organization's NISO member contact is, please contact us at nisohq@niso.org. There is no need to register separately. Check your institutional membership status here.
Location
Educational events are online programs. NISO uses the Zoom platform for the purpose of broadcasting our live events. Zoom provides apps for a variety of computing devices (tablets, laptops, etc.) To view the broadcast, you need a device that supports the Zoom app. Attendees may also choose to listen just to audio on their phones. Sign-on credentials include the necessary dial-in numbers, if that is your preference. Once notified of their availability, recordings may be viewed from the Zoom platform.
Registrants received sign-on instructions prior to the virtual event. If you have any questions, please contact NISO headquarters for assistance via email (nisohq@niso.org).