We need a new name for the fairly painful process of trying to tease meaning out of people’s unstructured HTML.
In any case downloading a bunch of stuff from JSTOR was not scraping. And he absolutely had authorised access to that data. They took exception to the quantity, mainly.
Also that he was planning, or already was, sharing the data for free. They made an example out of him because he believed such data should be for everybody.
So scraping means any harvesting of data now?
We need a new name for the fairly painful process of trying to tease meaning out of people’s unstructured HTML.
In any case downloading a bunch of stuff from JSTOR was not scraping. And he absolutely had authorised access to that data. They took exception to the quantity, mainly.
Scraping has multiple meanings.
Web scraping is a specific type of scraping, but data via APIs or even torrents could be considered a scrape, even if that data is nicely structured.
The commonality between them is they all have the implication that:
Any access patterns that broadly correspond to this could be considered scraping.
Sure. Fine. It means all of that now.
But what are we going to call the difficult thing that we have to do to coax, say, unstructured event listings into reasonably structured data?
If you want to refer specifically to web scraping then call it “web scraping” or “site scraping” or “HTML scraping”
Not difficult to get your meaning across.
Also that he was planning, or already was, sharing the data for free. They made an example out of him because he believed such data should be for everybody.