Discussion summary

A discussion on home DNA sequencing highlights the importance of error correction, with some users sharing personal experiences and skepticism about its practical value.

What the discussion says

  • Error correction in sequencing is crucial, especially with low coverage.
  • Some users find the process dense but manageable with AI tools.
  • Skepticism exists about the practical benefits of personal genome data.
10x coverage genuinely washes out errors because they're closer to independent.
joel_liu
Uploading protocols to ChatGPT makes understanding easier.
bmwoolf

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Bringing everything to your doorstep and everything at your feet and everything near your fingertips is just what all industries are trying to accomplice. The cartoon Animation Wall-E has scenes in it where they show obese humans doing everything through a screen though notice their legs and feet and it's as if they've mutated to a point where they aren't able to walk anymore and all their transport is through a hovering chair cum bed.
  • I watched wall-e in theater, and when that scene came on, i remember muttering 'what bs'; since then, i recall that scene every time i see a situation of 'convenience at all costs'; metaphorically they were pretty accurate even after discounting ozempic influence;
  • How can I know if the the results I get are real or just some garbage?
  • Run it multiple times, and then get professionals in on it as well, and then compare.
  • via comparison to other nucleotide sequences. This is called sequence alignment: https://en.wikipedia.org/wiki/Sequence_alignment
  • As someone with experience (albeit almost 20 years out of date) experience of wet lab DNA collection and sequencing - this stuff is hard, you will fail a lot, and you will fail a lot more if you don't have an extremely clean environment to do this in. And once you have data, you should be asking yourself how accurate it is given the environment you collected it in, you should be looking at correlated sequence errors that are not taken into account[1].

    But also: genetic counselling is a real thing that real people study. Please don't ask an LLM questions about what your genes are going to do to you without having access to someone who has the ability to contexualise the data and put you in touch with relevant experts. I have a PhD in this and I would not trust myself to be able to interpret data about myself in a detached and rational way.

    (And: why is the link to Molecular Biology of the Cell to the 6th edition, when the 7th came out 4 years ago? Random fact: the first three editions were co-authored by my supervisor during my first PhD attempt, who went on to demonstrate that Roger Penrose's ideas about the importance of microtubules in chemotaxis in E. coli were absolute bullshit. Great guy)

    [1] I spent a while analysing very early (by commercial standards) Illumina data in 2007, and being able to align stuff to reference genomes made it possible to identify certain biases. Nanopore technology is likely to have more of those, and if you don't have the ability to take those into account you may have a very bad time

  • > this stuff is hard, you will fail a lot, and you will fail a lot more if you don't have an extremely clean environment to do this in

    The Oxford Nanopore sequencing technology is one of the most robust to use. You need to buy some kit - but defo doable. Nothing compared pouring your own gel and doing radioactively labelled Sanger reactions :-)

    Though you could just go to an sequencing company that services labs ( that just does sequencing outsourcing - rather than a personal genome company ).

    Totally agree on the dangers around interpretation.

  • Nanopore data is a lot easier to analyze than short read sequencing data. You just don’t get the same alignment/assembly issues: these things sequence incredibly long reads.

    (also a biochemist, MSc)

  • I like the privacy conscious aspects. Apart from the obvious issue of "run it through Claude" how many of those referenced analysis tools are entirely open source or at least run locally? Would have liked to see that in the article.
  • At a quick glance, they all seem to have published their source code and they do run locally.
  • Reminds me of the Gloing Plant Project. I never got my glowing flower but would have settled for the instruction manual, also never created :(

    https://en.wikipedia.org/wiki/Glowing_Plant_project

    By the by, can't seen to bring up the actual site linked on this post.

  • You can get one here, you'll have to wait until next year though: https://light.bio/
  • I want to sequence, but I absolutely do not want any company or government or church to have access to my data. When the author says:

    > "I have a VCF, I can run it through tools like VEP, ClinVar, gnomAD, PharmGKB (highly recommend), Gene Inspector, or Claude"

    I am assuming my data is now within the hands of some of the very entities I do not want to have access to my data ... true?

  • Untrue in the cases of VEP, ClinVar, gnomAD. These are either offline, open-source tools or databases. You can download or query these and no one would be the wiser.

    Claude on the other hand - yeah you are giving your data away. But that step really isn't necessary.

    By running a standard pipeline you could get a VCF (File containing the Variants in your genome) and each variant would be annotated. You can check all the annotated genes and figure out if these variants are pathogenic, likely pathogenic, likely not pathogenic or benign.

  • This is so cool. Thanks for doing this. The fact that we have this in a palm sized object is just crazy. Also, if/when we have a similar sized device for doing CRISPR .... umm i should stop here - it's becoming the plot of Gattaca
  • https://www.the-odin.com/whole-genome-sequencing-30x/

    If you want it quick and cheap(er) - 599.00

  • A service is not the same as the equipment
    by j45
  • If it's an US-based lab, aren't they subject to CLIA with all its retention requirements?

    For $7.5k+ you get a guaranteed privacy (as other comments suggest, other properties may vary, but at least the data never leaves your home).

  • I wish this had some discussion of the results. The earlier reports about this sensor and process were very mixed. It’s a cool process either way, but I’d like to know how usable the real world output can be.
  • I've bee thinking about starting a company where I fish roots out of your sewer and identify the plant (by sequence if necessary) that you have to kill so your sewer doesn't collapse as soon as it otherwise would.

    $100 to stave off that $10000 sewer replacement for a few years would be worth it to a lot of people

  • Do it!
  • Hard to get a plumber to knock on my door for under $100. Maybe you mean $500? Or is it $100 for the lab bit only?
  • How will you reach out to enough people so they are aware and can order the service? I'm thinking about the minimum viable business opportunity here.
  • That’s a really clever idea, I would definitely pay for that in the right circumstances.

    Now that I think about it - could you just pour some sort of biodegradable broad-spectrum herbicide down the drain to get the same effect for cheaper?

  • I don't really understand this. How much better are things for killing a specific plant than whatever you would use to kill a superset of the plants that might be there? Or treating with a few different plant species killers. And wouldn't those options cost less than the cost of this service?
    by mhb
  • https://www.envirodna.com/

    https://www.naturemetrics.com/species-detection

    https://www.ednacollab.org/industry/

    https://wilderlab.co/

    These companies focus on environmental DNA - some are more on the level of local government monitoring, some are for private customers.

  • > This is intended to be read by AI- please just copy and paste the URL of this and have ChatGPT walk you through it. If you have AR glasses, even better, since the AI can walk you through the whole protocol.

    What kind of magic is going on here, am I missing something?

  • I feel it's actually kind of smart. Most people won't be reading the blog post themselves, they'd ask GPT to understand the text and fetch the summary or whatever is relevant to them. The author has directly made the resource such that it is optimised for the output after that mostly-everyone-would-do-this step.
  • Hi, author here- I intended it to be hands-free so you can upload to ChatGPT/Claude and talk to it. I found it easier to follow the protocol each time by talking to AI rather than having to read from the computer every time I had to check something, reducing context-switching

    You can still read it, though it is pretty dense

  • I suspect the intention is to give specific but dense notes with minimal explanation, on the theory that the LLM will fill in the appropriate hand-holding along the way
    by alwa