Nearline Web Archiving

dc.contributor.authorXie, Zhiwuen
dc.contributor.authorNayyar, Kratien
dc.contributor.authorFox, Edward A.en
dc.date.accessioned2016-06-28T20:46:06Zen
dc.date.available2016-06-28T20:46:06Zen
dc.date.issued2016-06-23en
dc.description.abstractIn this paper, we propose a modified approach to realĀ­time transactional web archiving. It leverages the web caching infrastructure that is already prevalent on web servers. Instead of archiving web content at HTTP transaction time, in our approach the archiving happens when the cached copy expires and is about to be expunged. Before the deletion, all expired cache copies are combined and then sent to the web archive in small batches. Since the cache is purged at much lower frequency than HTTP transactions, the archival workload is also much lower than that for transactional archiving. To further decrease the processing load at the origin server, archival copy deduplication is carried out at the archive instead of at the origin server. It is crucial to note that the cache purging process is separate from those that serve the HTTP requests. It can be, and usually is set to lower priority. The archiving therefore occurs only when the server is not busy fulfilling its more mission critical tasks; this is much less disruptive to the origin server. This approach, however, does not guarantee that the freshest copy is archived, although the cache purging policy may be adjusted to attempt to bound the freshness of the archive.en
dc.identifier.urihttp://hdl.handle.net/10919/71648en
dc.language.isoen_USen
dc.relation.ispartof3rd International Workshop on Web Archiving and Digital Libraries (WADL2016)en
dc.rightsCreative Commons Attribution-ShareAlike 3.0 United Statesen
dc.rights.urihttp://creativecommons.org/licenses/by-sa/3.0/us/en
dc.subjectWeb archivingen
dc.subjectNearline web archivingen
dc.subjectApache web serveren
dc.subjectWeb cacheen
dc.titleNearline Web Archivingen
dc.typeArticleen

Files

Original bundle
Now showing 1 - 4 of 4
Loading...
Thumbnail Image
Name:
2016-WADL-nearline.pdf
Size:
91.48 KB
Format:
Adobe Portable Document Format
Description:
Submitted version
Loading...
Thumbnail Image
Name:
2016-WADL-nearline-slides.pdf
Size:
780.71 KB
Format:
Adobe Portable Document Format
Description:
Slides for presentation at WADL 2016
Name:
ApachewithWARC.mp4
Size:
36.92 MB
Format:
MP4 Container format for video files
Description:
Demonstration video
Name:
ApachewithWARC.webm
Size:
4.57 MB
Format:
The webm video container format
Description:
License bundle
Now showing 1 - 1 of 1
Name:
license.txt
Size:
1.5 KB
Format:
Item-specific license agreed upon to submission
Description: