Show simple item record

dc.contributor.authorXie, Zhiwuen
dc.contributor.authorNayyar, Kratien
dc.contributor.authorFox, Edward A.en
dc.description.abstractIn this paper, we propose a modified approach to real­time transactional web archiving. It leverages the web caching infrastructure that is already prevalent on web servers. Instead of archiving web content at HTTP transaction time, in our approach the archiving happens when the cached copy expires and is about to be expunged. Before the deletion, all expired cache copies are combined and then sent to the web archive in small batches. Since the cache is purged at much lower frequency than HTTP transactions, the archival workload is also much lower than that for transactional archiving. To further decrease the processing load at the origin server, archival copy deduplication is carried out at the archive instead of at the origin server. It is crucial to note that the cache purging process is separate from those that serve the HTTP requests. It can be, and usually is set to lower priority. The archiving therefore occurs only when the server is not busy fulfilling its more mission critical tasks; this is much less disruptive to the origin server. This approach, however, does not guarantee that the freshest copy is archived, although the cache purging policy may be adjusted to attempt to bound the freshness of the archive.en
dc.relation.ispartof3rd International Workshop on Web Archiving and Digital Libraries (WADL2016)en
dc.rightsCreative Commons Attribution-ShareAlike 3.0 United Statesen
dc.subjectWeb archivingen
dc.subjectNearline web archivingen
dc.subjectApache web serveren
dc.subjectWeb cacheen
dc.titleNearline Web Archivingen

Files in this item


This item appears in the following Collection(s)

Show simple item record

Creative Commons Attribution-ShareAlike 3.0 United States
License: Creative Commons Attribution-ShareAlike 3.0 United States