Web Search Engines: Part 1
Problems for Web Searchers:
1. Infrastructure: must have the computers and hardware to meet the number of demands in a given period of time.
2. Crawling algorithms: Bots that go around the internet and index it. Crawlers start with a list or queue of good 'seed URLs'- sites that have lots of links to to other good websites. They then add all the unseen links to the queue, and save the content for indexing. The keep doing this until they hit the end of the queue.
Speed: One crawler can't do the whole internet! Need multiple crawlers who are assigned to different URLs to work in parallel (hashing function). Each crawler machine has internal parallelism as well, with multiple threads working at once.
Politeness: Don't harass the website's servers!
Excluded content: must look at robot.txt to determine what content should not be crawled.
Duplicate content: Avoid it.
Continuous crawling: Have a priority queue so important URL's are checked more frequently than low value/static URL's.
Spam: Prevent it! Blacklists, etc.
Web Search Engines: Part 2
Indexing algorithms: Scans document for indexable terms. These are then ranked in terms of position and repetition to give importance.
Real indexers:
Scaling up: divide the load among many machines, and fill up memory space with partial inverted files, and then combine the partials.
Term look up: So many phrases, so little time. Engines use trees, hierarchies and 2 level structures to make things more efficient.
Compression: Save space, compress data structures. This also makes searches faster.
Phrases: produce lists of common phrases.
Anchor text: words used to describe a link. A strongly repeated anchor text gives a good clue as to what the website is about.
Link popularity: the more people link to you, the better you are. This is proof that life really is a popularity contest, no matter what your mother told you.
Query-independent score: high scores in other ways improve ranking, even if it doesn't match the query as well.
Query processing algorithms: simple query processor looks up each word in its dictionary and locates postings list. It scans the postings list for documents in common.
Make them faster! Skip unnecessary parts of the list, end the results list early, number the documents based on their decreasing query-independent scores. Another option: cache! Precompute and store HTML results pages for popular searches. Spit this out upon request.
Henzinger:
The first part of this article focused on the same issues related to web search engines as the previous two articles.
Where it differed was in the third section:
Content Quality: How do you deal with wrong or misleading information? This is a topic that has occupied librarians' attention for a while. We produce guides and tutorials and lists on how to filter out the 'junk' on the internet. You forget that search engines try to help out with that. It's not a question of tricking the search engine into giving results that are not appropriate to the query, but one of whether or not the information provided is correct, even if it answers the query. Page rank and hits are a good measure, but not perfect. Anchor text might be useful, but junky websites can still have quality links. The most plausible is text based analysis.
Quality Evaluation: Measure the number of clicks a given result gets, and then the number of click throughs from that website.
Web Conventions: These habits of websites must be adhered to for the search engine to be able to use them correctly.
Anchor Text: Text in the links describes the link.
Appropriate and interesting links for the website's audience, related to the website's content.
Meta tags: like metadata in a library catalog, meta tags in a webpage can describe the site's content.
Duplicate hosts: multiple domain names resolve at the same end site for increased visibility. This is why typing in "pubmed.gov" sends you to http://www.ncbi.nlm.nih.gov/sites/entrez/. This is called a mirror. Search engines run the risk of providing results for each of those names, even though they have identical content. However, it could be hard to tell that if the ads on the pages are slightly different from one viewing to the next. If a webcrawl of the full site is not complete, then it might appear that they are not duplicates. A good way to avoid them is to predict whether similar domain names are likely to be duphosts.
Vaguely-structured data: Prose on a website that is marked up with HTML to affect how it is seen by the viewer. This HTML can give clues to the website. Large text followed by small text can imply that the small text is further details about the large text. Pages with an image in the upper left are often personal pages. Pages with more meta mistakes are likely to be of lower quality.
These readings gave me some interesting insight into how search engines work. I know that none of them are specific, because the actual algorithms are tightly guarded secrets. But they do give a clue as to why we get the results that we do, and just how hard the programmers work to fight off the spammers and such.
Friday, October 10, 2008
Friday, October 3, 2008
Week 6: Preservation in Digital Libraries
Research Challenges in Digital Libraries
We must research digital libraries in order to get a grasp on where we can take them. They are too widespread and heterogenous to really understand anything that's going on at the moment. We also need to figure out how to preserve the digital libraries as they are now for future study.
Big Issues:
1. We must figure out how to deal with all the digital libraries and preserve them while using humans as infrequently as possible.
2. We must protect the digital archives now. They require a lot of effort to maintain, so we must find a way to do that while, again, using humans as infrequently as possible.
3. We need to look at economic and business models of digital libraries to see how we can maintain these things in ways beyond technology. How can we afford to keep them up?
4. In order to expand the usefulness of digital libraries, new technologies need to be created. This needs to happen in order to make DL's cheaper while using humans as infrequently as possible.
5. We need shared and scalable infrastructure to support digital libraries. Sequestering them within institutions prevents interoperability and scalability, which hinders the usefulness of digital libraries.
Open Archival Information System Reference Model: An Introductory Guide
Open: reference model was developed in an open public forum: anyone could participate.
Archival Information System: people and institutions who agree to preserve info and make it available.
An OAIS must:
1. Get the appropriate information.
2. Make sure they have long term control of the information.
3. Know their user community.
4. Have appropriate metadata for the user to understand the info.
5. Make sure information is totally preserved.
6. Make it available to the user.
Tasks of OAIS:
-Ingestion (of data)
-Preservation Planning
-Data Management
-Archival Storage
-Administration
-Access (of data to user)
Types of information packages:
-Submission Information Package
-Archival Information Package
-Disseminated Information Package
This model provides a formula for digital library producers to follow. By doing so, they could produce an efficient, effective digital library. The paper does not provide any guidance on the technology or infrastructure to make this happen, but it does provide the guideposts of what sorts of things the technology and infrastructure must do.
Preservation Management of Digitized Materials
- The authors state that guidance is needed for digital preservation. It seems to be a recurring theme.
We must research digital libraries in order to get a grasp on where we can take them. They are too widespread and heterogenous to really understand anything that's going on at the moment. We also need to figure out how to preserve the digital libraries as they are now for future study.
Big Issues:
1. We must figure out how to deal with all the digital libraries and preserve them while using humans as infrequently as possible.
2. We must protect the digital archives now. They require a lot of effort to maintain, so we must find a way to do that while, again, using humans as infrequently as possible.
3. We need to look at economic and business models of digital libraries to see how we can maintain these things in ways beyond technology. How can we afford to keep them up?
4. In order to expand the usefulness of digital libraries, new technologies need to be created. This needs to happen in order to make DL's cheaper while using humans as infrequently as possible.
5. We need shared and scalable infrastructure to support digital libraries. Sequestering them within institutions prevents interoperability and scalability, which hinders the usefulness of digital libraries.
Open Archival Information System Reference Model: An Introductory Guide
Open: reference model was developed in an open public forum: anyone could participate.
Archival Information System: people and institutions who agree to preserve info and make it available.
An OAIS must:
1. Get the appropriate information.
2. Make sure they have long term control of the information.
3. Know their user community.
4. Have appropriate metadata for the user to understand the info.
5. Make sure information is totally preserved.
6. Make it available to the user.
Tasks of OAIS:
-Ingestion (of data)
-Preservation Planning
-Data Management
-Archival Storage
-Administration
-Access (of data to user)
Types of information packages:
-Submission Information Package
-Archival Information Package
-Disseminated Information Package
This model provides a formula for digital library producers to follow. By doing so, they could produce an efficient, effective digital library. The paper does not provide any guidance on the technology or infrastructure to make this happen, but it does provide the guideposts of what sorts of things the technology and infrastructure must do.
Preservation Management of Digitized Materials
- The authors state that guidance is needed for digital preservation. It seems to be a recurring theme.
This book is to extensive to takes notes in much detail. However, it is an extremely interesting, useful guide for a novice in digital libraries to get a handle on the field. It introduces the reader to the vocabulary, provides reasons on why this information is vital, and explains how digital libraries are made, who uses them, what the rules and requirements are, and provides models for institutions to follow as they delve into this realm. Since this is a very new world, and many librarians are long out of library school, having this sort of resource, perhaps with additional instruction, they can get up to speed. Staying abreast of technological developments is important, and digital libraries are a huge part of that.
The National Digital Newspaper Project is an effort to "Chronicle America" by digitally preserving printed newspapers. It "also has a digital repository component that houses the digitized newspapers, supporting access and facilitating long-term preservation. Taking on access and preservation in a single system was both a deliberate decision and a deviation from past practices at LC." They wrote this paper to discuss the work done so far. Specifically, they discuss the preservation threats encountered by the project in 2 years.
Types of failures:
Media- Failure in the portable hard drives transporting the digital images from the awardees to LC. Fixed using 'fixity checks' as part of the transfer process and keeping a copy at the awardees until it was verified that LC had received it.
Hardware- Internal hard drives failed. They avoided data loss by using multiple HD arrays in a RAID 5 array with a hot spare. This prevented data loss in case one failed. Data was only lost when a second event occurred in the array while the system was rebuilding the harddrive using the hot spare.
Software- Three software problems occurred. The first involved a validation problem: records were put into the NDNP repository that had passed validation but 'did not conform to the appropriate NDNP profile'. This was fixed with new validation rules. The second was more problematic. During transformation, the newspaper title record had stripped the original METS record of the XML, and also, was producing invalid METS records. This broke the application, and also made parts of the data unreadable. The third problem occurred when the XFS file system was corrupted. This caused data loss. In a large, complex system such as this, it is harder to prevent problems, and to diagnose them when they occur. This is a serious flaw of huge digital libraries.
Operator- One error occurred when a series of files were deleted accidentally. Another occurred when the operator accidentally ingested the same batches multiple times, or perhaps did not purge a successful ingest before re-ingesting it. Many duplicates were produced.
The conclusions of the paper are that in a huge task such as this, errors are going to occur in many different ways, no matter what one does to protect against them. This makes performing a large digitization project extremely daunting, since one of the tasks is to make sure that the files are not only accessible but also permanently preserved.
This is Katie's favorite person in the world. His name is Kevin. Yes, all 3 of us have K names. It was not planned: Katie came prenamed, and we didn't have any choice in our names.
Labels:
Digital library,
preservation,
week 6,
weekly response
Muddiest Point 4
In XML, it seems like there are multiple ways of structuring things to get the same result. Are these rules hard and fast, or fairly soft?
Sunday, September 28, 2008
Flickr Assignment
To see some cool pictures of Batman comicbook covers, check out this URL:
http://www.flickr.com/photos/30893186@N05/
http://www.flickr.com/photos/30893186@N05/
Friday, September 26, 2008
Muddiest Point 3
Week 5: XML Galore!
"Introducing the Extensible Markup Language"
XML is extensible: it can be altered and added to indefinitely to tweak the language to suit the needs of the user. This makes it a robust language to use for digital libraries. As things change, XML can accommodate the changes without requiring a total overhaul of the system. Libraries like things that work that way, because it doesn't require them to reinvent the wheel. It is also useful for metadata, because the tags can be used for labeling different types of metadata.
"A Survey of XML Standards" is a good reference source for the different versions of XML because it provides other resources to look at for further instruction. The sheer number of versions illustrates the extensibility of XML.
"Extending your Markup" is an interesting and short overview of how XML works. Again, it is a good resource for a novice to look at to get started in this new world.
Major definitions:
DTD: document type definitions. This tags a given field as including a given type of information, such as author. They define the structure of the XML document.
DTD elements:
Nonterminal: they have a series of other choices or sequeneces. A DTD defining a book has sequences following it such as author.
Terminal: They do not have choices. They may include things like PC data, or are empty, or labeled as 'any'.
DTD attributes: do not prescribe order on the DTD, but include further information
Namespaces: to prevent conflict between two fields that use the same tag but in different contexts (email address vs. postal address) namespaces define the two as distinct. Do not play well with DTDs.
Linking: Goes beyond HTML to describe different types of linking
Xlink: describes how 2 documents can be linked
Xpointer: links 2 parts of the same document.
XPath: (used by Xpointer) describes the linking path
XSLT: Extensible Style Sheet Language Transformer: goes from XSL to HTML.
XML Schema: Overcome the limitations of DTDs (expression limited and non XML syntax)
Document definition markup language (DDML): define datatypes
Document content description (DCD)
Schema for object-oriented XML (SOX)
XML-Data (replaced by DCD)
"Introduction to XML schema"
Schema replace DTDs! They do the same things like define the element, define child elements, define the order of the elements, and other similar things. However, they are more extensible, richer and powerful, they support data types and namespaces and they are still XML. Essentially, they perform the same function as DTD's only better.
XML is extensible: it can be altered and added to indefinitely to tweak the language to suit the needs of the user. This makes it a robust language to use for digital libraries. As things change, XML can accommodate the changes without requiring a total overhaul of the system. Libraries like things that work that way, because it doesn't require them to reinvent the wheel. It is also useful for metadata, because the tags can be used for labeling different types of metadata.
"A Survey of XML Standards" is a good reference source for the different versions of XML because it provides other resources to look at for further instruction. The sheer number of versions illustrates the extensibility of XML.
"Extending your Markup" is an interesting and short overview of how XML works. Again, it is a good resource for a novice to look at to get started in this new world.
Major definitions:
DTD: document type definitions. This tags a given field as including a given type of information, such as author. They define the structure of the XML document.
DTD elements:
Nonterminal: they have a series of other choices or sequeneces. A DTD defining a book has sequences following it such as author.
Terminal: They do not have choices. They may include things like PC data, or are empty, or labeled as 'any'.
DTD attributes: do not prescribe order on the DTD, but include further information
Namespaces: to prevent conflict between two fields that use the same tag but in different contexts (email address vs. postal address) namespaces define the two as distinct. Do not play well with DTDs.
Linking: Goes beyond HTML to describe different types of linking
Xlink: describes how 2 documents can be linked
Xpointer: links 2 parts of the same document.
XPath: (used by Xpointer) describes the linking path
XSLT: Extensible Style Sheet Language Transformer: goes from XSL to HTML.
XML Schema: Overcome the limitations of DTDs (expression limited and non XML syntax)
Document definition markup language (DDML): define datatypes
Document content description (DCD)
Schema for object-oriented XML (SOX)
XML-Data (replaced by DCD)
"Introduction to XML schema"
Schema replace DTDs! They do the same things like define the element, define child elements, define the order of the elements, and other similar things. However, they are more extensible, richer and powerful, they support data types and namespaces and they are still XML. Essentially, they perform the same function as DTD's only better.
Friday, September 19, 2008
Muddiest Point week 3
Who assigns a DOI? Is it the creator of the digital object, or an outside organization?
Subscribe to:
Posts (Atom)
