A free database of university web links: data collection issues
This paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make avail...
| Autor: | |
|---|---|
| Tipo de documento: | artigo |
| Estado: | Versão publicada |
| Data de publicação: | 2002 |
| País: | España |
| Recursos: | Consejo Superior de Investigaciones Científicas (CSIC) |
| Repositório: | DIGITAL.CSIC. Repositorio Institucional del CSIC |
| OAI Identifier: | oai:digital.csic.es:10261/174381 |
| Acesso em linha: | http://hdl.handle.net/10261/174381 |
| Access Level: | Acceso aberto |
| Palavra-chave: | Web links Web Impact Factor Search engines Web crawler |
| id |
ES_7ec320d1205a85d4fe75e2721062f9ad |
|---|---|
| oai_identifier_str |
oai:digital.csic.es:10261/174381 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
A free database of university web links: data collection issuesThelwall, MikeWeb linksWeb Impact FactorSearch enginesWeb crawlerThis paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make available raw data for research that is not reliant upon the opaque techniques of commercial search engines. Basic tools for querying are also provided. The key issues concerning running an accurate web crawler are also discussed. Access is also given to the normally hidden crawler stop list with the aim of making the crawl process more transparent. The necessity of having such a list is discussed, with the conclusion that fully automatic crawling is not socially or empirically desirable because of the existence of database-generated areas of the web and the proliferation of the phenomenon of mirroringPeer reviewedEditorial CSIC201920192002info:eu-repo/semantics/articlehttp://purl.org/coar/resource_type/c_6501Publisher's versioninfo:eu-repo/semantics/publishedVersionhttp://hdl.handle.net/10261/174381reponame:DIGITAL.CSIC. Repositorio Institucional del CSICinstname:Consejo Superior de Investigaciones Científicas (CSIC)InglésNoinfo:eu-repo/semantics/openAccessoai:digital.csic.es:10261/1743812026-05-22T06:33:51Z |
| dc.title.none.fl_str_mv |
A free database of university web links: data collection issues |
| title |
A free database of university web links: data collection issues |
| spellingShingle |
A free database of university web links: data collection issues Thelwall, Mike Web links Web Impact Factor Search engines Web crawler |
| title_short |
A free database of university web links: data collection issues |
| title_full |
A free database of university web links: data collection issues |
| title_fullStr |
A free database of university web links: data collection issues |
| title_full_unstemmed |
A free database of university web links: data collection issues |
| title_sort |
A free database of university web links: data collection issues |
| dc.creator.none.fl_str_mv |
Thelwall, Mike |
| author |
Thelwall, Mike |
| author_facet |
Thelwall, Mike |
| author_role |
author |
| dc.subject.none.fl_str_mv |
Web links Web Impact Factor Search engines Web crawler |
| topic |
Web links Web Impact Factor Search engines Web crawler |
| description |
This paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make available raw data for research that is not reliant upon the opaque techniques of commercial search engines. Basic tools for querying are also provided. The key issues concerning running an accurate web crawler are also discussed. Access is also given to the normally hidden crawler stop list with the aim of making the crawl process more transparent. The necessity of having such a list is discussed, with the conclusion that fully automatic crawling is not socially or empirically desirable because of the existence of database-generated areas of the web and the proliferation of the phenomenon of mirroring |
| publishDate |
2002 |
| dc.date.none.fl_str_mv |
2002 2019 2019 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article http://purl.org/coar/resource_type/c_6501 Publisher's version info:eu-repo/semantics/publishedVersion |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10261/174381 |
| url |
http://hdl.handle.net/10261/174381 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
No |
| dc.rights.none.fl_str_mv |
info:eu-repo/semantics/openAccess |
| eu_rights_str_mv |
openAccess |
| dc.publisher.none.fl_str_mv |
Editorial CSIC |
| publisher.none.fl_str_mv |
Editorial CSIC |
| dc.source.none.fl_str_mv |
reponame:DIGITAL.CSIC. Repositorio Institucional del CSIC instname:Consejo Superior de Investigaciones Científicas (CSIC) |
| instname_str |
Consejo Superior de Investigaciones Científicas (CSIC) |
| reponame_str |
DIGITAL.CSIC. Repositorio Institucional del CSIC |
| collection |
DIGITAL.CSIC. Repositorio Institucional del CSIC |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869411774122426368 |
| score |
15,198674 |