Discussion summary

A genomics guide tailored for engineers received positive feedback, with suggestions to expand on genome-scale analysis and RNA stability. Some commenters noted its basic level and potential for future articles.

What the discussion says

  • Some users suggest adding sections on whole-genome analysis and polygenic traits.
  • Others recommend including topics like RNA degradation and downstream analysis.
  • There is appreciation for the guide's relevance to computer scientists and engineers.
“This is super cool, thanks”
— egyptianblue
“Would be great to see more on genome-scale analysis”
— murzynalbinos

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > In plants and animals, DNA is broken up into a number of large sequences called chromosomes that are tucked into the nucleus.

    This is a weird description, because ... it is not really "broken up". Each chromosome could be shuffled and put into different cells in different numbers. Now, it is unlikely that the resulting cell would be viable or useful, but my contention here is the "broken up" part. Chromosomes are just a way to handle the genome set. There are reasons why bacteria do not have chromosomes and this has mostly to do with the amount of DNA. To call this "breaking up" is a very strange description. (Size is not the only reason; duplication of the DNA before cell division is another important factor; bacteria usually have just one origin of replication, eukaryotes have several on each chromosome, otherwise the S-phase in the cell cycle would simply take too long.)

    > Each genome is a biochemical database that, if properly accessed, can inform how our bodies function.

    This is also a very strange description, aka "biochemical database". Not everything in a genome has a role with regards to biochemistry or metabolism. Some is just regulatory RNA; some of this relates to metabolism, but you also have e. g. piwiRNA or silencers of transposons and so forth. That in itself has only very rarely a biochemical function, with some exceptions (e. g. I would classify tRNA as related to metabolism, and many viruses have tRNA or use tRNA as quick-starters, but most of those regulatory RNAs do not have any function for metabolism directly, other than e. g. repurposing energy towards their own reproduction).

    To me it seems as if the article was written by an engineer. That's fine, but it also means that the thinking is quite biased. Genetics is not quite so easy to engineer; a good example are leaky promoters used in synthetic biology (just ask the people who use such promoters how to make them un-leaky) or off-target cleavage effects in CRISPR-Cas(9 or whatever is used); I am pretty certain they'll give excuses as to why 100% accurate gene therapy isn't yet ready for the masses. And they'll do that for quite some years to come, I bet, usually hiding behind "it will cost too much" - when in reality, it should cost very little, if it were to work, rather than this just becoming the new meta-milking scheme.

  • This is very very nice. when you are reading this, just keep this in the back of your mind - inside a cell- things are floating around constantly at a very high speed. those things do not have any crisp shape or boundary. so how do we tell them apart? they are phase separated. if you put an oil drop in water, you can still see the oil drop and water and tell them apart. that's a very high degree of phase separation. inside a cell the degree of phase separation is much lower. just putting this out here so that you could appreciate the complexity of the biology that you are reading. my wife educated me on this a bit.
  • This website was made by St. Jude Children's Research Hospital. AFAIK they are a non-profit which runs treatment clinical trials on children cancer patients and doesn't bill them for anything.

    The course was mentioned in a recent whoishiring thread. Sounds like it could be a purposeful place to work.

    I'm not affiliated in any way, just found it interesting.

    by yreg
  • I've worked for a year in a lab doing cancer genomics and had to learn everything from scratch, since my background is in computer science.

    It's definitely possible to learn enough to be productive within a few months, but to actually comprehend and understand the underlying biology takes much, much longer. I still don't understand much of what is presented by people from other labs outside of my specialty.

  • One part that people from the software side tend to underestimate is how fuzzy and analog everything in biology is. Genomics look more predictable and organized at first, but even these parts are quite fuzzy and subject to all kinds of physical effects.

    I'd strongly recommend in reading up on the parts of cell biology that come after this. Otherwise you'll get the wrong impression of how messy biology actually is.

  • If you're an engineer and want to go deeper into the core algorithms behind genomics, there's a book / course called Bioinformatics Algorithms. It was a punishing read when I was going through it a few years ago (but rewarding). It's probably much better now given the state of AI.

    [1] https://cogniterra.org/course/64/info

  • As a software developer who spent nearly 10 years founding and building a genomic startup, this is a good start, but does have a lot of vast oversimplifications and a few inaccuracies. The people making this know what they're doing, so I'm sure these are known shortcomings they likely deemed necessary for a quick introduction. You'd need a further study at the end to start being able to do some real-world work.
  • This guide is also made from me (or some of the me from a couple years back). I haven't read the whole thing yet and it's probably clearly stated at some point (though one can deduce it with the beginning already) but the surprise for me was that this field is highly statistical. Before starting I had the (very) naive view that it was possible to read the genome as one reads a file and look at what's going on. But the sequencing technics (and accompanying algorithms) only allow to statistically read the genome. So variants/mutations found are only found with a given statistical certainty. If the sample wasn't well prepared for example it could be that this certainty is ultimately not high enough to do a proper analysis/diagnostic. It's a fascinating field (try to watch a video on sequencing by expansion, to feel how sci-fi this field actually is) that is very hard to approach with only high-school biology level and this guide is really well done to sort of bridge this first gap.

Explore Birbla archives