A free database of university web links: data collection issues

This paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make avail...

ver descrição completa

Detalhes bibliográficos
Autor: Thelwall, Mike
Tipo de documento: artigo
Estado:Versão publicada
Data de publicação:2002
País:España
Recursos:Consejo Superior de Investigaciones Científicas (CSIC)
Repositório:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/174381
Acesso em linha:http://hdl.handle.net/10261/174381
Access Level:Acceso aberto
Palavra-chave:Web links
Web Impact Factor
Search engines
Web crawler
id ES_7ec320d1205a85d4fe75e2721062f9ad
oai_identifier_str oai:digital.csic.es:10261/174381
network_acronym_str ES
network_name_str España
repository_id_str
spelling A free database of university web links: data collection issuesThelwall, MikeWeb linksWeb Impact FactorSearch enginesWeb crawlerThis paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make available raw data for research that is not reliant upon the opaque techniques of commercial search engines. Basic tools for querying are also provided. The key issues concerning running an accurate web crawler are also discussed. Access is also given to the normally hidden crawler stop list with the aim of making the crawl process more transparent. The necessity of having such a list is discussed, with the conclusion that fully automatic crawling is not socially or empirically desirable because of the existence of database-generated areas of the web and the proliferation of the phenomenon of mirroringPeer reviewedEditorial CSIC201920192002info:eu-repo/semantics/articlehttp://purl.org/coar/resource_type/c_6501Publisher's versioninfo:eu-repo/semantics/publishedVersionhttp://hdl.handle.net/10261/174381reponame:DIGITAL.CSIC. Repositorio Institucional del CSICinstname:Consejo Superior de Investigaciones Científicas (CSIC)InglésNoinfo:eu-repo/semantics/openAccessoai:digital.csic.es:10261/1743812026-05-22T06:33:51Z
dc.title.none.fl_str_mv A free database of university web links: data collection issues
title A free database of university web links: data collection issues
spellingShingle A free database of university web links: data collection issues
Thelwall, Mike
Web links
Web Impact Factor
Search engines
Web crawler
title_short A free database of university web links: data collection issues
title_full A free database of university web links: data collection issues
title_fullStr A free database of university web links: data collection issues
title_full_unstemmed A free database of university web links: data collection issues
title_sort A free database of university web links: data collection issues
dc.creator.none.fl_str_mv Thelwall, Mike
author Thelwall, Mike
author_facet Thelwall, Mike
author_role author
dc.subject.none.fl_str_mv Web links
Web Impact Factor
Search engines
Web crawler
topic Web links
Web Impact Factor
Search engines
Web crawler
description This paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make available raw data for research that is not reliant upon the opaque techniques of commercial search engines. Basic tools for querying are also provided. The key issues concerning running an accurate web crawler are also discussed. Access is also given to the normally hidden crawler stop list with the aim of making the crawl process more transparent. The necessity of having such a list is discussed, with the conclusion that fully automatic crawling is not socially or empirically desirable because of the existence of database-generated areas of the web and the proliferation of the phenomenon of mirroring
publishDate 2002
dc.date.none.fl_str_mv 2002
2019
2019
dc.type.none.fl_str_mv info:eu-repo/semantics/article
http://purl.org/coar/resource_type/c_6501
Publisher's version
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10261/174381
url http://hdl.handle.net/10261/174381
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv No
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
eu_rights_str_mv openAccess
dc.publisher.none.fl_str_mv Editorial CSIC
publisher.none.fl_str_mv Editorial CSIC
dc.source.none.fl_str_mv reponame:DIGITAL.CSIC. Repositorio Institucional del CSIC
instname:Consejo Superior de Investigaciones Científicas (CSIC)
instname_str Consejo Superior de Investigaciones Científicas (CSIC)
reponame_str DIGITAL.CSIC. Repositorio Institucional del CSIC
collection DIGITAL.CSIC. Repositorio Institucional del CSIC
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869411774122426368
score 15,198674