Screen Scraping vs Web Scraping: What’s the Difference and Which One Do You Need?

Anyone who has spent any time near data collection will have come across both of these terms. They are often mentioned in the same sentence or seem to mean the same thing. In fact, they do not, and confusing them can cause a project to go off on the wrong track before it has even begun.

Here is a straightforward explanation of screen scraping and web scraping. Rather than using technical terms for the sake of it, I’m telling you how each method works, in what situations each is appropriate, and how you can decide between them.

Screen Scraping vs Web Scraping What's the Difference and Which One Do You Need

The Short Answer

Screen scraping involves reading the content that a program displays on the screen, while web scraping consists of reading the code and data that a website sends to your browser.

That does sound like a minor difference, yet it does affect the tools you use, the speed at which the job runs, and how frequently it breaks.

What Screen Scraping Actually Is

Consider the way in which you obtain information from an application—you open it, look at the window, and then locate the number or name that you want. Screen scraping is software that carries out that same process for you.

The program examines the way another application displays information and extracts the values that a person would read. It could carry out this task in a number of ways:

  • Reading text from a terminal. Older business systems often show plain text at fixed positions, and a script can read those positions.
  • Reading the elements of a window. Desktop operating systems expose buttons, fields, and labels so that assistive tools like screen readers can use them. Automation software can read the same information.
  • Reading an image. When nothing else is available, the program takes a screenshot and uses optical character recognition (OCR) to turn the picture of the text back into text.

What is common to all three is that the program does not interact with the database or with the source of the data; it only deals with the surface.

Which is why screen scraping is generally resorted to only when there is no other means of access, for example, with old software that has no API or with a document that exists only as a scan.

What Web Scraping Actually Is

A web page isn’t a picture; there is an HTML document behind what you see in the browser, and usually a stream of structured data that the page loads in the background.

Web scraping proceeds directly to that level; the script requests the page just as a browser would, obtains the code, and then selects the elements it needs: the product name, the price, the headline, and the list of links.

Since the script is analyzing the structure rather than the appearance, it doesn’t matter what font the site uses or where a box is positioned on the page; the only thing that matters is that the price is within a specific tag or field.

Certain websites create their content within the browser using JavaScript, so the initial code received is essentially empty. In such cases, people make use of a headless browser, which is a genuine browser that runs without displaying a window. This type of browser loads the page completely, and then the script looks at the result.

The Differences That Matter

The way they compare in terms of the things you actually experience in a project is as follows.

Where the data comes from. Screen scraping can work on almost any application, because every application shows something. Web scraping only works on websites and the data they serve.

Speed. Speed is achieved because screen scraping has to wait until a screen has been drawn before it can read it; web scraping, on the other hand, bypasses the drawing process and retrieves the data directly, as a result of which it is much faster.

Accuracy. When you read a value from structured code, you get an exact result, but when you read it from pixels, you are making an estimate. Moreover, OCR in particular can mistake one character for another, for example confusing a zero with a capital O.

How it breaks. A screen scraper may fail if a layout changes, a font is altered, or a window is resized. In fact, it might continue to run and give you the incorrect value without informing you. A web scraper fails when the site changes the code, and this kind of change is generally easier to detect.

Scale. Since screen scraping is linked with sessions and rendering, it doesn’t scale easily. Web scraping is capable of dealing with a large number of pages, which is the reason it is the usual choice when collecting public web data.

For a fuller side-by-side comparison, including the techniques behind each approach, this guide on screen scraping vs web scraping is worth a read.

Where People Get Confused

The main source of confusion is the middle ground.

Is the act of a script reading the output when a headless browser renders a page considered screen scraping or web scraping? One could justify seeing it either way. Although the page has been rendered as if it were being viewed on a screen, the script is still reading the document, not an image of it. Generally, people refer to this as web scraping, and I believe that is the appropriate term.

Another source of confusion is habit. Many people tend to use the expression ‘screen scraping’ as a general term to describe any form of automated data collection. Whenever a colleague or a client uses that term, it is advisable to find out exactly what they mean before proceeding with any plans.

How to Choose

A simple rule works well here: use the most structured source you can get.

  1. The best approach is to use the official API or the data export.
  2. When the data is available on a public website, scrape the HTML or the data itself.
  3. When the content does not appear on the page until after JavaScript has been executed, use a headless browser.
  4. When the data is contained in a desktop application, and there is no API, use UI automation.
  5. When the data is available only in the form of an image or a scan, then use OCR.

With each step downwards on the list, you lose a certain amount of speed and a certain amount of reliability; therefore, you go down only when the step above it is not available.

A Few Practical Notes

Check your results. Make sure that you check your results since this is especially important in the case of screen scraping and OCR. Ensure that there is a method in place for verifying key values, as silent errors represent the actual danger.

Expect maintenance. Maintenance should be expected since applications and websites are subject to change; regardless of the method you choose, you should make arrangements for the day when it stops working, and someone has to carry out the repair.

Respect the rules. Just because a method is technically feasible doesn’t mean that it is allowed. The terms of websites, software licenses, copyright rules,s and privacy regulations all apply and can vary from one place to another. Limit yourself to the data which you are permitted to collect and never use anyone else’s login information without their permission.

Think about volume on the web side. If you collect public web data in large amounts, many requests from one IP address will often get slowed down or blocked. That is the reason web scraping setups commonly use rotating proxies. Screen scraping of desktop or internal systems does not need them, because that work stays inside your own network.

Conclusion

There are two methods available that enable software to read information that was intended for people, the difference being that they read it from different locations. Screen scraping involves looking at the surface of an application, while web scraping consists of examining the code and data beneath a web page.

So the choice between screen scraping vs .web scraping usually makes itself once you know where your data lives. If it is on a website, web scraping is faster, more accurate, and easier to scale. If it is locked inside an old system or an image, screen scraping may be the only door available, and that is exactly what it is for.

Popular on OTW Right Now!

Add a Comment

Your email address will not be published. Required fields are marked *