Sign in

Travis Reid

@treid803.bsky.social
24 followers 33 following 34 posts

PhD Student at ODU CS and member of @webscidl.bsky.social.

PostsRepliesMedia
Reposted by Travis Reid
himarshaj.bsky.social @himarshaj.bsky.social · 31/08/2026
Excited to share that our team received an IMLS award for CiteCast, a project to identify and preserve web resources referenced in U.S. TV news! @WebSciDL ✍️: ws-dl.blogspot.com/2026/08/2026... Drs. @acnwala.bsky.social, @ibnesayeed.bsky.social, @phonedudemln.bsky.social, & @weiglemc.bsky.social
ws-dl.blogspot.com
2026-08-31: IMLS Grant Awarded on Identifying and Preserving Web Resources Referenced in Archived U.S. Local and National Television News
The Web Science and Digital Libraries Research Group at Old Dominion University.
124
Reposted by Travis Reid
Michael L. Nelson @phonedudemln.bsky.social · 22/06/2026
2026 WS-DL Research Expo Tuesday, June 23, 2pm EDT All are welcome: members, alumni, friends, etc. github.com/oduwsdl/2026... 5 students, 1 from each @webscidl.bsky.social prof, will give a ~15 min summary of their current research. Status updates from profs & alumni too! (zoom link in repo)
github.com
GitHub - oduwsdl/2026-research-expo: 2026 ODU Web Science and Digital Libraries (WSDL) Research Expo
2026 ODU Web Science and Digital Libraries (WSDL) Research Expo - oduwsdl/2026-research-expo
139
Reposted by Travis Reid
himarshaj.bsky.social @himarshaj.bsky.social · 09/06/2026
Excited to share that I passed my PhD proposal defense on April 24, 2026 -- another step closer to finishing my PhD 🎓 Special thanks to my dissertation committee: @weiglemc.bsky.social, @phonedudemln.bsky.social, @acnwala.bsky.social, & Erika Frydenlund for their guidance & support.
slideshare.net
Client Challenge
154
Reposted by Travis Reid
Michael L. Nelson @phonedudemln.bsky.social · 27/05/2026
Recently, Sawood (ibnesayeed.bsky.social) & I uncovered a tricky web archiving replay problem. An API call several layers of JS deep had superfluous browser size args that would intermittently load the wrong JSON for the archived HTML ws-dl.blogspot.com/2026/05/2026... (all images/links are SFW)
ws-dl.blogspot.com
2026-05-26: URL Arguments in API Calls Can Cause Intermittent Temporal Violations While Replaying Archived Web Pages
The Web Science and Digital Libraries Research Group at Old Dominion University.
135
Reposted by Travis Reid
Michael L. Nelson @phonedudemln.bsky.social · 14/02/2026
Hussam Hallak @hussamhallak.bsky.social of @webscidl.bsky.social describes using Google Sheets to do a bulk upload of URLs to the "Save Page Now" service of the Wayback Machine @waybackmachine.bsky.social ws-dl.blogspot.com/2026/02/2026...
ws-dl.blogspot.com
2026-02-12: How to archive web pages in bulk using the Internet Archive Google Sheets service
The Web Science and Digital Libraries Research Group at Old Dominion University.
048
Travis Reid @treid803.bsky.social · 13/02/2026
They also described how Browsertrix assists with quality assurance by allowing users to rate the quality of archived web pages and by computing 3 metrics, which involve comparing extracted text, counts of HTTP status codes for resources, and screenshots during crawl and replay.
000
Travis Reid @treid803.bsky.social · 13/02/2026
In this paper, they discussed the features of multiple @webrecorder.net tools. Browsertrix is their web archiving platform that uses Browsertrix Crawler for archiving web pages, ArchiveWeb.page for patching web pages, & ReplayWeb.page to replay archived resources.
100
Travis Reid @treid803.bsky.social · 13/02/2026
For this blog post, I summarized @bitarchivist.net, @wilkinson.graphics, and @ilya.webrecorder.net's paper, “High Fidelity Web Archiving of News Sites and New Media with Browsertrix.” @webscidl.bsky.social ws-dl.blogspot.com/2026/02/2026...
ws-dl.blogspot.com
2026-02-13: Paper Summary: "High Fidelity Web Archiving of News Sites and New Media with Browsertrix"
The Web Science and Digital Libraries Research Group at Old Dominion University.
144
Reposted by Travis Reid
Lesley Frew @lesleyelisabeth.bsky.social · 03/02/2026
In "Detecting and reconstructing trustworthy edit histories using web archives", I talk about different markers of edit history on webpages and how web archives can be used to verify or refute them. ws-dl.blogspot.com/2026/02/2026... @webscidl.bsky.social
035
Reposted by Travis Reid
Michael L. Nelson @phonedudemln.bsky.social · 23/01/2026
Travis Reid @treid803.bsky.social @webscidl.bsky.social summarizes Brenda Reyes Ayala's @brendareyesayala.bsky.social 2025 paper: "Towards a better QA process: Automatic detection of quality problems in archived websites using visual comparisons" ws-dl.blogspot.com/2026/01/2026...
ws-dl.blogspot.com
2026-01-22: Paper Summary: "Towards a better QA process: Automatic detection of quality problems in archived websites using visual comparisons"
The Web Science and Digital Libraries Research Group at Old Dominion University.
036
Travis Reid @treid803.bsky.social · 10/12/2025
5) Successful replay of ads loaded in iframes with the src attribute of "about:blank" depended upon a given browser's service worker implementation. A Chromium bug stopped service workers from accessing resources inside of this type of iframe, which prevented replay. (9/9)
000
Travis Reid @treid803.bsky.social · 10/12/2025
4) When loading Flashtalking web page ads outside of ad iframes, the ad script requested a non-existent URL, which prevented the replay of ad resources. (8/9)
100
Travis Reid @treid803.bsky.social · 10/12/2025
3) During crawling and replay sessions, Google's and Amazon's ad scripts generated URLs with different random values, because the random number generator’s seed is not the same during the crawl and replay sessions. This prevented archived ads' replay. (7/9)
100
Travis Reid @treid803.bsky.social · 10/12/2025
2) During 2023, Brozzler was incompatible with versions of Chrome released after version 111.0.5563.110, which prevented ads from being archived. This incompatibility was resolved during 2024. This thread describes the problem: x.com/TReid803/sta... (6/9)
x.com
Travis Reid on X: "A month ago I noticed that @internetarchive's browser-based crawler (Brozzler) was not working with the most recent stable version of Google Chrome when using the brozzle-page command. GitHub Issue: https://t.co/q1EfUsEicH #WebArchiveWednesday @WebSciDL (1/4)" / X
A month ago I noticed that @internetarchive's browser-based crawler (Brozzler) was not working with the most recent stable version of Google Chrome when using the brozzle-page command. GitHub Issue: https://t.co/q1EfUsEicH #WebArchiveWednesday @WebSciDL (1/4)
100
Travis Reid @treid803.bsky.social · 10/12/2025
1) Before August 2023, Internet Archive's Save Page Now (SPN) excluded ad services' ads & URLs with ad related file and directory names. In April 2025, a new option, "Disable ad blocker," was introduced in SPN, allowing logged-in users to archive ads. (5/9)
100
Travis Reid @treid803.bsky.social · 10/12/2025
Five problems with archiving & replaying ads during 2023: 1. IA's Save Page Now excluded ads 2. Brozzler's incompatibility with Chrome 3. Google & Amazon ad URLs with random values 4. Flashtalking ads requested unarchived URL 5. Replay of ads differed depending on browser (4/9)
100
Travis Reid @treid803.bsky.social · 10/12/2025
We created a dataset of 279 ads by archiving 17 web pages from SimilarWeb’s top websites worldwide. Dataset of 279 archived ads: github.com/savingads/Re... We also created a web page to display ads from our dataset: savingads.github.io/themed_ad_co... (3/9)
100
Travis Reid @treid803.bsky.social · 10/12/2025
We also published a blog post on Information Matters, which provides a summary of our process for creating a dataset of 279 archived web ads and the problems we identified while archiving and replaying these ads. (2/9) #WebArchiveWednesday @webscidl.bsky.social informationmatters.org/2025/12/prob...
informationmatters.org
Problems With Archiving and Replaying Web Advertisements - Information Matters
Advertisements are an integral part of our cultural heritage, and this extends to online web advertisements. Unlike print ads, web ads are dynamic and interactive, which makes them difficult to archiv...
102
Travis Reid @treid803.bsky.social · 10/12/2025
Our JASIST article, "Problems with archiving and replaying current web advertisements", was recently published: onlinelibrary.wiley.com/share/author... @wiley.com Co-authors: Alex Poole, Hyung Wook Choi, Christopher Rauch, @machawk1.bsky.social, @phonedudemln.bsky.social, @weiglemc.bsky.social (1/9)
onlinelibrary.wiley.com
114
Travis Reid @treid803.bsky.social · 12/06/2025
Goel et al.’s crawler (Jawa) removes some non-deterministic JavaScript code so that the replay of an archived web page does not change if different users replay it. Removing this code resulted in storage savings of 41% and improved the crawling throughput by 39%.
000
Travis Reid @treid803.bsky.social · 12/06/2025
While working on the Saving Ads project, we encountered similar problems that involved JavaScript code generating URLs with random values that differed during crawl time and replay time. A Google SafeFrame URL is an example where a random value in the URL caused a replay problem.
100
Travis Reid @treid803.bsky.social · 12/06/2025
When non-determinism causes variance in resources’ URLs it results in failed requests, which prevents resources from loading. Goel et al. matched a requested URL with a crawled URL by removing the query string (querystrip) and using Levenshtein distance (fuzzy matching).
100
Travis Reid @treid803.bsky.social · 12/06/2025
Sources of non-determinism identified by Goel et al: > Server-side state > Client side state > Client characteristics > JavaScript's Date, Random, and Performance APIs
100
Travis Reid @treid803.bsky.social · 12/06/2025
In this blog post, I summarized Goel et al.’s "Jawa: Web Archival in the Era of JavaScript". They identified sources of non-determinism that cause replay problems and created a crawler that removes non-deterministic code. #WebArchiveWednesday @webscidl.bsky.social ws-dl.blogspot.com/2025/06/2025...
ws-dl.blogspot.com
2025-06-11: Paper Summary: "Jawa: Web Archival in the Era of JavaScript" (Goel et al. OSDI '22)
The Web Science and Digital Libraries Research Group at Old Dominion University.
114
Reposted by Travis Reid
himarshaj.bsky.social @himarshaj.bsky.social · 13/05/2025
✍️Our @webscidl.bsky.social trip report for ACM #CAPWIC2025 is now live! Check out the highlights by Yasasi, Lesley, Kritika, & I: ws-dl.blogspot.com/2025/05/2025... @weiglemc.bsky.social, @phonedudemln.bsky.social, @openmaze.bsky.social, @nirdslab.bsky.social
ws-dl.blogspot.com
2025-05-12: ACM Capital Region Celebration of Women in Computing (CAPWIC) 2025 Trip Report
The Web Science and Digital Libraries Research Group at Old Dominion University.
014
Reposted by Travis Reid
Michele Weigle @weiglemc.bsky.social · 08/05/2025
This semester, all of my federal grants were terminated. In a new blog post, I write about the importance of federal funding for academic research and the impact of these cuts. @WebSciDL #NEH #IMLS #NSF ws-dl.blogspot.com/2025/05/2025...
ws-dl.blogspot.com
2025-05-08: Thank you to NEH, IMLS, DoD Minerva, and NSF
The Web Science and Digital Libraries Research Group at Old Dominion University.
11411
Reposted by Travis Reid
Michael L. Nelson @phonedudemln.bsky.social · 25/04/2025
I appreciate the chance to present "A Vision for Trustworthy Web Archiving" at the DPC Members Forum and Networking Event - Americas meeting at Vanderbilt University. bit.ly/Nelson-DPC2025
bit.ly
Nelson -- A Vision for Trustworthy Web Archiving -- DPC Members Forum and Networking Event - Americas 2025-04-24
A Vision for Trustworthy Web Archiving Michael L. Nelson @phonedudemln.bsky.social with: Sawood Alam, Vikas Ashok, Mohamed Aturban, John Berlin, Justin Brunelle, Kritika Garg, Hussam Hallak, Himarsha ...
136
Travis Reid @treid803.bsky.social · 12/02/2025
Our tech report provides additional details not covered in the blog posts: bsky.app/profile/trei... (7/7)
001
Travis Reid @treid803.bsky.social · 12/02/2025
To learn more about the replay problems identified while creating this dataset you can read this blog post: ws-dl.blogspot.com/2024/12/2024... (6/7)
ws-dl.blogspot.com
2024-12-04: Problems With Replaying Ads That Use iframes
The Web Science and Digital Libraries Research Group at Old Dominion University.
100
Travis Reid @treid803.bsky.social · 12/02/2025
We also created a web page that allows us to view all of the information from the dataset including the replay of the archived ads: savingads.github.io/themed_ad_co... (5/7)
100
Travis Reid @treid803.bsky.social · 12/02/2025
To identify ads that were not able to replay in the containing web page that loaded the ads during the crawling session, we used ReplayWeb.page’s URL search feature (replayweb.page/docs/user-guide/exploring/) and our Display Archived Ads tool (github.com/savingads/Display-Archived-Ads). (4/7)
100
Travis Reid @treid803.bsky.social · 12/02/2025
When archiving these web pages, we utilized: > Web archiving services >> Internet Archive's Save Page Now >> Arquivo.pt >> archive.today >> Conifer > Browser-based tools >> ArchiveWeb.page >> Browsertrix Crawler >> Brozzler (3/7)
100
Travis Reid @treid803.bsky.social · 12/02/2025
Our dataset of 279 ads was created by archiving 17 web pages from SimilarWeb's top websites worldwide. Dataset: github.com/savingads/Re... (2/7)
github.com
100
Travis Reid @treid803.bsky.social · 12/02/2025
One of the goals for the Saving Ads project was to create a dataset of advertisements from the live web. This blog post describes our process for creating this dataset: ws-dl.blogspot.com/2025/02/2025... #WebArchiveWednesday @webscidl.bsky.social (1/7)
ws-dl.blogspot.com
2025-02-10: Creating a Dataset of Archived Web Ads
The Web Science and Digital Libraries Research Group at Old Dominion University.
102
Travis Reid @treid803.bsky.social · 06/02/2025
We were able to create a dataset of 279 ads by archiving 17 web pages from SimilarWeb’s top websites worldwide. Dataset of 279 archived ads: github.com/savingads/Re... We also created a web page to display ads from our dataset: savingads.github.io/themed_ad_co... (10/10)
012
Travis Reid @treid803.bsky.social · 06/02/2025
To identify ads that were not able to replay in the containing web page that loaded the ads during the crawling session, we used ReplayWeb.page’s URL search feature (replayweb.page/docs/user-guide/exploring) and our Display Archived Ads tool (github.com/savingads/Display-Archived-Ads). (9/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
5) Successful replay of ads loaded in iframes with the src attribute of "about:blank" depended upon a given browser's service worker implementation. A Chromium bug stopped service workers from accessing resources inside of this type of iframe, which prevented replay. (8/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
4) When loading Flashtalking web page ads outside of ad iframes, the ad script requested a non-existent URL, which prevented the replay of ad resources. (7/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
We created an example web page that used ad code from Google’s pubads_impl_2023020201.js script to determine how the random values were generated for a Google SafeFrame. Demo web page for generating random numbers and Google SafeFrames: treid003.github.io/random_Value... (6/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
3) During crawling and replay sessions, Google's and Amazon's ad scripts generated URLs with different random values, because the random number generator’s seed is not the same during the crawl and replay sessions. This prevented archived ads' replay. (5/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
2) During 2023, Brozzler was incompatible with versions of Chrome released after version 111.0.5563.110, which prevented ads from being archived. This thread describes the problem: x.com/TReid803/sta... (4/10)
x.com
x.com
100
Travis Reid @treid803.bsky.social · 06/02/2025
1) Before August 2023, Internet Archive's Save Page Now (SPN) excluded ad services' ads & URLs with ad related file and directory names. After August 2023, SPN still excluded ads loaded on a web page & only allowed ad resources if the user directly archived the ad's URL(s) (3/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
Five problems with archiving & replaying ads during 2023: 1. IA's Save Page Now excluded ads 2. Brozzler's incompatibility with Chrome 3. Google & Amazon ad URLs with random values 4. Flashtalking ads requested unarchived URL 5. Replay of ads differed depending on browser (2/10)
100
Travis Reid @treid803.bsky.social · 06/02/2025
Our new tech report (arxiv.org/abs/2502.01525): "Archiving and Replaying Current Web Advertisements…" describes the archiving and replay problems we encountered while creating a dataset of 279 archived ads. #WebArchiveWednesday @webscidl.bsky.social @machawk1 @phonedudemln @weiglemc (1/10)
235