←back to thread

646 points blendergeek | 4 comments | | HN request time: 0s | source
Show context
quchen ◴[] No.42725651[source]
Unless this concept becomes a mass phenomenon with many implementations, isn’t this pretty easy to filter out? And furthermore, since this antagonizes billion-dollar companies that can spin up teams doing nothing but browse Github and HN for software like this to prevent polluting their datalakes, I wonder whether this is a very efficient approach.
replies(9): >>42725708 #>>42725957 #>>42725983 #>>42726183 #>>42726352 #>>42726426 #>>42727567 #>>42728923 #>>42730108 #
1. grajaganDev ◴[] No.42725708[source]
I am not sure. How would crawlers filter this?
replies(2): >>42725835 #>>42726294 #
2. captainmuon ◴[] No.42725835[source]
Check if the response time, the length of the "main text", or other indicators are in the lowest few percentile -> send to the heap for manual review.

Does the inferred "topic" of the domain match the topic of the individual pages? If not -> manual review. And there are many more indicators.

Hire a bunch of student jobbers, have them search github for tarpits, and let them write middleware to detect those.

If you are doing broad crawling, you already need to do this kind of thing anyway.

replies(1): >>42727490 #
3. marginalia_nu ◴[] No.42726294[source]
You limit the crawl time or number of requests per domain for all domains, and set the limit proportional to how important the domain is.

There's a ton of these types of of things online, you can't e.g. exhaustively crawl every wikipedia mirror someone's put online.

4. dylan604 ◴[] No.42727490[source]
> Hire a bunch of student jobbers,

Do people still do this, or do they just off shore the task?