As I write, I am busy on another computer programing an Arduino board to make little lights flash on and off. Thee guy next to me has made his play Billy Jean...at double speed, which is kind of annoying and fun at the same time.
Arduino is interesting. You have the little circuit board which you wire up, and then you connect it too the computer using a USB, write a program (heh), get it to run, and, if you're lucky, a little light flashes. Or Billy Jean plays at double speed.
It's so much fun!
To put this in a little bit of context, I'm in the middle of a two week synthetic-biology course. People keep trying to get me to do programing, which is slightly disturbing. I am enjoying playing with Arduino though. Almost as much as I enjoyed constructing the bed-side tables last night :D
Field of Science
-
-
Change of address1 year ago in Variety of Life
-
Change of address1 year ago in Catalogue of Organisms
-
-
Earth Day: Pogo and our responsibility1 year ago in Doc Madhattan
-
What I Read 20241 year ago in Angry by Choice
-
I've moved to Substack. Come join me there.1 year ago in Genomics, Medicine, and Pseudoscience
-
-
-
-
Histological Evidence of Trauma in Dicynodont Tusks7 years ago in Chinleana
-
Posted: July 21, 2018 at 03:03PM8 years ago in Field Notes
-
Why doesn't all the GTA get taken up?8 years ago in RRResearch
-
-
Harnessing innate immunity to cure HIV10 years ago in Rule of 6ix
-
What kind of woman would pray for health or use spiritual healing?10 years ago in Epiphenom
-
-
-
-
-
-
post doc job opportunity on ribosome biochemistry!11 years ago in Protein Evolution and Other Musings
-
-
Blogging Microbes- Communicating Microbiology to Netizens12 years ago in Memoirs of a Defective Brain
-
Re-Blog: June Was 6th Warmest Globally12 years ago in The View from a Microbiologist
-
-
-
The Lure of the Obscure? Guest Post by Frank Stahl14 years ago in Sex, Genes & Evolution
-
-
Lab Rat Moving House15 years ago in Life of a Lab Rat
-
Goodbye FoS, thanks for all the laughs15 years ago in Disease Prone
-
-
Slideshow of NASA's Stardust-NExT Mission Comet Tempel 1 Flyby15 years ago in The Large Picture Blog
-
in The Biology Files
Showing posts with label computers. Show all posts
Showing posts with label computers. Show all posts
Yes but what does it do...
I am currently trying to get myself to finished writing an essay (rather terrifyingly my first essay of term) on the different approaches to gene annotation in vertebrates. As I've just woken up (afternoon naps seem like such a good idea until you wake up with a mouth that feels like a hamster died in it) I thought I'd give a quick summary of gene annotation methods:
Gene annotation is the 'interesting' bit of genomics. Quite a lot of gene sequencing work has been done, some of it (especially the human bits) very highly publicised. And while genome sequencing is probably useful (more on that maybe in a more ethically-inclined post) on it's own it's not terribly exciting. You're left with a big database full of mindless streams of nucleotides and one bit embarrassing question:
What does it all do?
Gene annotation attempts to answer that; trying to work out which proteins each gene codes for, essentially what the end function of the genome is, what each piece of DNA is used for. There are two main methods: just using DNA, and using data from protein/cDNA sources. Both of these methods can be either comparative or non-comparative:
1) Just using DNA: Non-Comparative
This relies on getting a program such as GENSCAN to, quite literally, scan along the DNA looking for the beginning and end of genes based on sequence patterns it had been told to recognise. Not so good for function, but useful enough for finding the damn genes in the first place. Also relatively cheap and you can go run it overnight.
2)Just using DNA: Comparative
Like it says, this compares your DNA with other previously annotated pieces of DNA to see if there are any very similar bits it can ascribe function to. It's a good starting point, especially now the pool of annotated genomes is increasing, but it's really bad at finding gene start point, especially when there are 'introns', or bits of DNA that are not actually turned into protein. Which is around 95% of the human genome incidentally. (an e.g of this, if anyones interested, is TWINSCAN)
3) cDNA/Protein data: Non-comparitive
cDNA, just to clarify, is DNA that has been reverse transcribed from RNA templates; i.e itt's all the DNA that will get turned into protein, and without any of the introns. A good way to use this is to make cDNA 'libraries' i.e all the cDNA within the cell stored on plasmids, choose one at random, see what it makes and, at the same time, find where it is in the genome. Simple and useful.
4) cDNA/Protein data: Comparative
This compares your genome with bits of cDNA from other genomes, where the cDNA has known function. Protein comparison is even more useful as seeing what protein your protein most resembles provides structural information, as well as functional and allows you to build up homologous families of proteins with similar function (if you have enough genomes). Also if you have enough protein data you can say you're doing 'proteomics' and the more 'omics' words in your project, the more funding you're likely to get :)
By the way, all of these comparative methods are based on homologous evolutionary relationships between the genomes, so anyone who says that scientists never use evolution is WRONG. (and probably pissing off the evodevo people as well)
As always, any questions are welcomed, leave them in the comments and I'll get back to you.
Gene annotation is the 'interesting' bit of genomics. Quite a lot of gene sequencing work has been done, some of it (especially the human bits) very highly publicised. And while genome sequencing is probably useful (more on that maybe in a more ethically-inclined post) on it's own it's not terribly exciting. You're left with a big database full of mindless streams of nucleotides and one bit embarrassing question:
What does it all do?
Gene annotation attempts to answer that; trying to work out which proteins each gene codes for, essentially what the end function of the genome is, what each piece of DNA is used for. There are two main methods: just using DNA, and using data from protein/cDNA sources. Both of these methods can be either comparative or non-comparative:
1) Just using DNA: Non-Comparative
This relies on getting a program such as GENSCAN to, quite literally, scan along the DNA looking for the beginning and end of genes based on sequence patterns it had been told to recognise. Not so good for function, but useful enough for finding the damn genes in the first place. Also relatively cheap and you can go run it overnight.
2)Just using DNA: Comparative
Like it says, this compares your DNA with other previously annotated pieces of DNA to see if there are any very similar bits it can ascribe function to. It's a good starting point, especially now the pool of annotated genomes is increasing, but it's really bad at finding gene start point, especially when there are 'introns', or bits of DNA that are not actually turned into protein. Which is around 95% of the human genome incidentally. (an e.g of this, if anyones interested, is TWINSCAN)
3) cDNA/Protein data: Non-comparitive
cDNA, just to clarify, is DNA that has been reverse transcribed from RNA templates; i.e itt's all the DNA that will get turned into protein, and without any of the introns. A good way to use this is to make cDNA 'libraries' i.e all the cDNA within the cell stored on plasmids, choose one at random, see what it makes and, at the same time, find where it is in the genome. Simple and useful.
4) cDNA/Protein data: Comparative
This compares your genome with bits of cDNA from other genomes, where the cDNA has known function. Protein comparison is even more useful as seeing what protein your protein most resembles provides structural information, as well as functional and allows you to build up homologous families of proteins with similar function (if you have enough genomes). Also if you have enough protein data you can say you're doing 'proteomics' and the more 'omics' words in your project, the more funding you're likely to get :)
By the way, all of these comparative methods are based on homologous evolutionary relationships between the genomes, so anyone who says that scientists never use evolution is WRONG. (and probably pissing off the evodevo people as well)
As always, any questions are welcomed, leave them in the comments and I'll get back to you.
Disclaimer: This post was written while half asleep. Any spelling/grammer mistakes are therefore completely the fault of the writers Brain On Sleep.
Speaking to computers
I am a Lab Rat. I work in laboratories, I try to figure out gels and results. I pipette tiny amounts of liquid into various locations. I try to see patterns and shapes and fit things together. This is what I do. I also occasionally venture into the world of Literature and try to find patterns there (but only as a hobby unfortunately).
What I don't do is computers. I can think of over twenty ways to re-phrase the instruction 'look for the comma' (probably over thirty ways if I'm allowed to use the 'synonyms' feature in word) but not one of those ways works for a computer.
Unfortunately there are some things that even Lab Rats need computers for. For instance, searching for a particular domain of DNA, pulling all the results out of a standard BLAST search, and then taking only the relevant information from that. To slightly clarify, BLAST is a bioinformatics programme with a huge database of information about every known and officially sequenced protein and DNA sequence. You type in your sequence (or your name, if you feel bored) and it shows you what proteins it matches on the database. Unfortunately it provides quite a large amount of information about each one, so this is where programming comes in, you tell the computer which bits of the database you want.
Now I did do computer science for IGCSE. I know what pseudocode is, and how to use it, and could probably stab a guess at setting commands up in the right sequence as well. Where I fall apart is the bit after that, translating the pseudocode into computer speak. There are some phrases I quite literally cannot do.
e.g:
IF comma is present
THEN stop
concatenate new information to old (concatenate=attach or add to)
OK. Fine. That works. But how do you say 'comma is present' in computer? The comma is not equal to anything, so that's out. Nor is the comma in relation to anything, it's just a comma in the middle of the string of writing. I have no idea at all how to tell the computer to find me a comma, Perl (which I am using to programme) seems to have no random squiggle that means 'find' or 'look for'.
Another thing that flummoxed me was the concatenation. All the online tutorials showed you exactly how to concatenate, but only if you already knew what the phrases were:
Tutorial: To concatenate x and y type x.=y
Lab Rat: I DON'T KNOW WHAT X AND Y ARE!!! In fact I'm looking for them because I don't know what they are. I want to find that out!;
Tutorial: *is no help at all*
Lab Rat: *Kicks computer, then hold up a large picture of a comma in front of the screen* Just find this, OK? See this picture, find something that looks like this and then give all the writing in front of it to me;
(the semicolon is computer for 'end of line'. I do not know why computers do this when almost every living person uses a full stop).
Computer: *Is not impressed*
I did actually get there in the end, to the surprise and delight of both myself and my supervisor (and probably the computer as well) I managed to get it to do sort of what I wanted. Unfortunately when we looked back over the raw data from our BLAST results we realised that there was a lot more information we wanted, so a lot more code has to be written. And here we hit another problem. The BLAST databases are truly amazing but just not particularly well organised. Some of them have the important information stored under /notes, while others have a separate field called /function. One item we saw even had the full protein function listed under /name. This means that to get all the information we need, we'll have to pull out each of these fields for every single protein, which will provide us with a lot of useless notes that we don't really need.
Lab Rat: Just give me the useful stuff, OK?;
Computer: Variable 'useful' not defined. Random computer squiggles, out of cheese error.
Lab Rat: *gives up on computers*;
Computer: *gives up on Lab Rat*
What I don't do is computers. I can think of over twenty ways to re-phrase the instruction 'look for the comma' (probably over thirty ways if I'm allowed to use the 'synonyms' feature in word) but not one of those ways works for a computer.
Unfortunately there are some things that even Lab Rats need computers for. For instance, searching for a particular domain of DNA, pulling all the results out of a standard BLAST search, and then taking only the relevant information from that. To slightly clarify, BLAST is a bioinformatics programme with a huge database of information about every known and officially sequenced protein and DNA sequence. You type in your sequence (or your name, if you feel bored) and it shows you what proteins it matches on the database. Unfortunately it provides quite a large amount of information about each one, so this is where programming comes in, you tell the computer which bits of the database you want.
Now I did do computer science for IGCSE. I know what pseudocode is, and how to use it, and could probably stab a guess at setting commands up in the right sequence as well. Where I fall apart is the bit after that, translating the pseudocode into computer speak. There are some phrases I quite literally cannot do.
e.g:
IF comma is present
THEN stop
concatenate new information to old (concatenate=attach or add to)
OK. Fine. That works. But how do you say 'comma is present' in computer? The comma is not equal to anything, so that's out. Nor is the comma in relation to anything, it's just a comma in the middle of the string of writing. I have no idea at all how to tell the computer to find me a comma, Perl (which I am using to programme) seems to have no random squiggle that means 'find' or 'look for'.
Another thing that flummoxed me was the concatenation. All the online tutorials showed you exactly how to concatenate, but only if you already knew what the phrases were:
Tutorial: To concatenate x and y type x.=y
Lab Rat: I DON'T KNOW WHAT X AND Y ARE!!! In fact I'm looking for them because I don't know what they are. I want to find that out!;
Tutorial: *is no help at all*
Lab Rat: *Kicks computer, then hold up a large picture of a comma in front of the screen* Just find this, OK? See this picture, find something that looks like this and then give all the writing in front of it to me;
(the semicolon is computer for 'end of line'. I do not know why computers do this when almost every living person uses a full stop).
Computer: *Is not impressed*
I did actually get there in the end, to the surprise and delight of both myself and my supervisor (and probably the computer as well) I managed to get it to do sort of what I wanted. Unfortunately when we looked back over the raw data from our BLAST results we realised that there was a lot more information we wanted, so a lot more code has to be written. And here we hit another problem. The BLAST databases are truly amazing but just not particularly well organised. Some of them have the important information stored under /notes, while others have a separate field called /function. One item we saw even had the full protein function listed under /name. This means that to get all the information we need, we'll have to pull out each of these fields for every single protein, which will provide us with a lot of useless notes that we don't really need.
Lab Rat: Just give me the useful stuff, OK?;
Computer: Variable 'useful' not defined. Random computer squiggles, out of cheese error.
Lab Rat: *gives up on computers*;
Computer: *gives up on Lab Rat*
Subscribe to:
Posts (Atom)