Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • One is reminded of Damien Hirst's famed retort to a critic who said "Well I could have pickled a shark" ... "But you didn't, did you. I did."
    by jmkd
  • Reminds me of Teddy Roosevelt's "Citizenship in a Republic" speech, where he talked about the man in the arena:

    > It is not the critic who counts; not the man who points out how the strong man stumbles or where the doer of deeds could have done them better. The credit belongs to the man who is actually in the arena, whose face is marred by dust and sweat and blood; who strives valiantly; who errs, and comes short again and again, because there is no effort without error and shortcoming; but who does actually strive to do the deeds; who knows the great enthusiasms, the great devotions; who spends himself in a worthy cause; who at the best knows in the end the triumph of high achievement, and who at the worst, if he fails, at least fails while daring greatly, so that his place shall never be with those cold and timid souls who know neither victory nor defeat.

    https://www.presidency.ucsb.edu/documents/address-the-sorbon...

  • > Sergey Brin and Larry Page came up with this precise algorithm, i.e., PageRank, which was one of the key algorithms that helped catapult Google into a household name and made them tons of money. Both Sergey and Larry were grad students at Stanford, so their coming up with such an amazing algorithm doesn’t seem surprising.

    No, Brin wasn’t a co-inventor of PageRank.

    https://patents.google.com/patent/US7058628B1/en

  • Inventor in the context of a U.S. patent has a very narrow and technical meaning, different from the broader concept of "coming up with something." Larry Page and Sergey Brin developed PageRank together as grad students in Stanford's Digital Library Project, led by Hector Garcia-Molina. This is simply undisputable and both are listed as co-authors of the PageRank paper [1]. The list of inventors on a patent, especially nowadays, should not be seen as a historical finding about credit.

    [1] https://www.semanticscholar.org/paper/The-PageRank-Citation-...

  • Yes anybody can be in the right place at the right time.
  • You could also have invented a wheel because it's so obvious.
  • The problem with the wheel is not inventing it so much as finding ways to use it. Aside from capstans and pulleys, the wheel is mostly only useful in combination with good infrastructure or complicated machinery.
    by kqr
  • Let's not forget that PageRank unintentionally spawned the SEO industry.
  • True
  • i don't think that's correct. SEO has been popular long before google. many websites were overloaded with long lists of unrelated keywords at the bottom.

    google became the most popular search engine exactly because those SEO keyword spam lists didn't work anymore and, after a while, they vanished. blogs, on the other hand, worked exceptionally well, because they used to link to other blogs' entries a lot, and blogs also used to be (comparably) high quality content.

    so the main SEO guideline for getting a good page rank used to be: have a blog/news and publish lots of high quality content that bloggers link to.

  • The blame should not rest at Pagerank's feet. The moment the web became a place to make money SEO would have cropped up to target any other ranking system, had that become the defacto standard. In fact the ranking system was not raw Pagerank's, it was one of the many signals used to rank a page. It was the resulting ranking that SEO targets, not Pagerank specifically.
  • As, I think, Page points out in the patent, PageRank's idea comes from Science Citation Index. That was an inverted list of scientific references, where you could look up an scientific paper in an expensive set of bound books and find all the papers in which it was later referenced. You can then use this to see who's getting referenced a lot, which is an ego trip in academia. Academic libraries had copies of that index. Now everybody has that kind of info, but when it had to be done by hand, it was hard.

    Inverting the huge, sparse matrix of references for PageRank was expensive. Originally, Google did it about once a week. The big breakthrough was when someone (who?) figured out how to do it incrementally at scale.

  • Thinking of an algorithm in the abstract is one thing, implementing it at scale is another.

    Yes you could have invented PageRank, but could you also have invented MapReduce, BigFiles/Google File System (GFS), Google Web Server, Bigtable, Protobuf? Then spun up fault-tolerant clusters consisting of cheap commodity PC hardware in an era where AWS wasn't even an idea yet? Then invented the concept of Borg to manage this hardware globally?

  • When Google started, Beowulf clusters were all the rage. So it wouldn't be a big conceptual leap at all. But productionizing it is serious work.
  • Well, I was a child in 1996, so probably not.

    Tying relevancy to link frequency was definitely a novel idea at the time, even if it seems "obvious" or simple in retrospect.

    by smcg
  • Here are two excellent videos that explain and visualize the PageRank algorithm:

    * [2020-06-17] Spanning Tree - "How Google's PageRank Algorithm Works" (5m16s): https://www.youtube.com/watch?v=meonLcN7LD4

    * [2022-05-23] Reducible - "PageRank: A Trillion Dollar Algorithm" (25m25s): https://www.youtube.com/watch?v=JGQe4kiPnrU

  • PageRank is fascinating, since it is so easy to explain.

    Yet, this is not even half the work. It like a third of the way.

    Before you could have invented PageRank, you must think in graphs. That is possible in 1996, but not as widespread as today.

    After you invented PageRank, you still need to deploy it. Again, possible but challenging as well. Is Python performant enough in 96? Can you afford more than 4MB RAM?

    At least Lego will not sue you for using their bricks to build a server rack in 1996.

  • > Before you could have invented PageRank, you must think in graphs.

    For anyone puzzling over what is the graph-based perspective, there's actually a very elegant and simple mathematical way to derive page rank.

    Define the directed graph of web links / citations through its adjacency matrix: web pages as nodes and links as directed edges. Represent a web surfer as a random walker on this graph, and run forth (simulate) the probability distribution of where they might end up.

    If you've studied linear algebra you would see an immediate analogy between the PageRank algorithm how the power method is used to find the dominant eigenvector (of the transition matrix, which is a normalized version of the adjacency matrix).

    This process is basically equivalent to implementing the diffusion process generated by the discrete graph laplacian. And simulating the random walk to find the long term stationary probability dstribution over nodes is akin to finding the zero eigenvector of this graph laplacian -- because it must generate "zero change" on the fixed point state.

  • > you must think in graphs. That is possible in 1996, but not as widespread as today.

    Could you elaborate? I'm a very graph-oriented thinker, and I was never aware this was some kind of declining skill (For context I was not alive in '96).

  • I actually built this, and shipped it, in 1996, with no knowledge of page rank, citation analysis, or bibliometrics, for an internal/external search engine for the Envirolink web site. Envirolink was a directory of environmental web sites, so they already had a list of URLs to crawl. The reason it was feasible to build was that it was a fairly constrained list of URLs, it wasn't the entire web.

    I didn't really know what I was doing (I was 17), but it was an awesome unpaid summer internship. There were two parts of the search engine - a crawler and the search engine. Both were written in Perl.