←back to thread

A graph explorer of the Epstein emails

(epstein-doc-explorer-1.onrender.com)
322 points cratermoon | 1 comments | | HN request time: 0s | source
Show context
pickpuck ◴[] No.45958412[source]
What if we extended this idea beyond one dataset to all discrete news events and entities: people, organizations, places.

Just like here you could get a timeline of key events, a graph of connected entities, links to original documents.

Newsrooms might already do this internally idk.

This code might work as a foundation. I love that it's RDF.

replies(10): >>45958506 #>>45958629 #>>45959158 #>>45959273 #>>45959323 #>>45959385 #>>45960015 #>>45960134 #>>45960357 #>>45963779 #
jandrewrogers ◴[] No.45959323[source]
This has been attempted many times. They all fail the same way.

These general data models start to become useful and interesting at around a trillion edges, give or take an order of magnitude. A mature graph model would be at least a few orders of magnitude larger, even if you aggressively curated what went into it. This is a simple consequence of the cardinality of the different kinds of entities that are included in most useful models.

No system described in open source can get anywhere close to even the base case of a trillion edges. They will suffer serious scaling and performance issues long before they get to that point. It is a famously non-trivial computer science problem and much of the serious R&D was not done in public historically.

This is why you only see toy or narrowly focused graph data models instead of a giant graph of All The Things. It would be cool to have something like this but that entails some hardcore deep tech R&D.

replies(5): >>45959382 #>>45960002 #>>45960262 #>>45960362 #>>45961019 #
theteapot ◴[] No.45960362[source]
> It would be cool to have something like this ..

Aren't LLMs something like this?

replies(1): >>45960459 #
1. djtango ◴[] No.45960459[source]
An LLM probabilistically produces tokens over its model which is why it can hallucinate whilst an actual graph model would not have that issue