GWAS 란?
다음을 요약해보면..
일단 GWAS는 genome wide association study의 약자. 뭔말인고 하니 genome wide 하게 association study를 한다는 의미. 이건 또 무슨 말인고 하니 genome wide 하게 DNA를 관측하고 이 DNA의 차이 혹은 변화를 질병과 같은 변수와 연관지여 연구를 한다는 의미.
사람의 genome은 서로 99.9%가 동일하다. 그래서 차이는 거의 없을 거고 이 차이가 나름 중요한 역할을 할거다(질병이라던지 관련해서). 그럼 이 개인간의 차이 즉 common SNP은 몇개나 있을까? 10 million쯤 있을거라 생각되고.. 그럼 GWAS를 연구할려면 이 10M SNP를 다 detecting 해야 하냐? 아니다. 과학자들이 HapMap project라는 것을 진행해서 SNP의 조합인 haplotype들을 찾아놨기 때문에 돈은 많이 싸졌다. 물론 요즘 NGS 값이 워낙 싸서 더 싸졌다.
그럼 어떤식으로 GWAS를 진행할까?
많은 다수 그러니까 수천명 규모에서 수만명까지 특정 질병군(case)와 일반군(control)의 사람이 필요하고(이때 confounding factor인 성, 종족등을 통일(?)시킨다)
1.각각의 사람을 genotyping 한다. 그러니까 SNP를 찾는다.
2.quality control of genotype data.
3.genotype과 phenotype간의 통계적으로 연관이 있는지 확인. chi-squared 같은 걸 이용
4.연관된 주변의 추가적인 SNP의 genotyp으로 fine mapping of association signal. 그니까 3번 step에서 유의하다고 생각되는 snp 주변의 SNP들을 추가적으로 genotyping 해서 haplotype을 경험적으로다가 만들어 낸다.
5.그 담에 다른 집단에다가 적용. 이땐 지정된 SNP로다가만 testing 한다.
6.biological validation
GWAS tools
많다. CRAN 의 task view의 genetics에서도 보면 GWAS를 할 수 있는 많은 package를 소개한다. 또한 broad institute 에서 만든 PLINK도 많이 사용한다.
Showing posts with label biological background. Show all posts
Showing posts with label biological background. Show all posts
Tuesday, October 11, 2011
Thursday, August 18, 2011
Wednesday, July 27, 2011
STR (short tandem repeat)
갑자기 뜬금없이 STR에 대해 조사하게 되었다. 원래 모르던 거니 알아보면서 정리 해보자.
STR이 뭐냐? 위키에 따르면 두개 이상의 nucleotide가 연속적으로 반복되어 있는것. 보통 non-coding intron에서 나타난단다. STRP (STR Polymorphism)은 사람마다 STR loci에 repeat의 수가 다름으로서 생기는 다형성. 법의학적으로 굉장히 많이 사용된단다.
여기까지만 하고 DB부터 검색해 보자.
구글부터 검색을 해보자면.. 여러개 나온다. ATCC, YHRD, ENFSI, STRBase. STRBase가 웹페이지는 허접해 보이나 왠지 좀 academic 해 보이는 관계로 이걸 위주로 보겠다.
STRBase의 DB논문은 2001년에 NAR에 나왔다. 이 DB는 NIST(National Institute of Standards and Technology)에서 관리하는 DB로 뭐 STR DNA marker, 관련 논문 정리, 관련 연구 기관, STR multiplex kits 등 STR 과 관련된 정보는 다 모아놓은 DB라고 할 수 있겠다. 여기까지가 abstract 내용이고
STRBase에서 제공하는 intro를 보자면 forensic DNA typing의 역사서 부터 시작해서 어떤 케이스에 STR이 사용되는지 실험kit, 기계, base call signal, 등등 아주 간략하게 그림만 나와있는데 마지막 슬라이드가 흥미롭다. CODIS(COmbined Dna Index System)라고 FBI의 DB가 있단다. 아래 그림에 나와있는 13 개의 dna loci 에 대해 법의학적으로 DNA profiling을 한다고 한다.
아 근데 이게 좀.. 음 STR이 microsatellite의 하나인 것 같은데.. 위키에서도 보면 두 개의 정의가 merge하도록 suggestion 되어 있고 다음 pdf를 봐도 microsatellite = STS + STR로 되어 있다. 음.. 나와 같은 고민을 한사람이 있었다. 여기 참조뭐.. 여튼
여기까지가 대충 한번 훑어 본거고. 그럼 내가 알아봐야 하는건 무엇인가. 이 STR의 길이의 variation과 maximum length. 음 근데 이거 STRBase에 잘 나와있다. 특히나 cat, dog, cattle의 경우 굉장히 편리하게 table로 정리해놨다. 음.. 사이즈가 대략 maximum 몇백 하는군..
하지만 내가 궁극적으로 찾아야 하는건 뭐냐. 식물에서의 STR. googling 하자마자 나온게 재밌게도 대마초의 STR에 관한 논문. STR이 법의학적으로 사용이 많이 되서 그런지 내가 찾아보면서 본 논문들은 거의 Forensic Science International 에서 나온 듯. abstract를 보자면 93개의 대마를 5개의 STR marker 로 profiling 했단다. 이 5 loci 에서 총 79 개의 allele이 확인됐다. 그담에 AMOVA로 genetic variation을 찾고 fibre crop accession과 drug accession의 variation이 어떻고 뭐하고 뭐한다는데.. 뭐래는 건지 모르겠다(결론은 drug종과 fibre종간의 차이를 확실하게 구분짓지는 못하나 drug 종간에는 사용할만하다 뭐 이런거 같은). 여튼 확인하고자 하는 건 STR의 maximum length. 아래 표가 그것을 의미하는게 아닐까 한다. 얼추 2~300 정도 하는거 같네.
그 다음에 드는 생각이 대마도 있는데 arabidopsis는 없을까? 역시 없을리 없다. 다음 논문을 보자(그닥 최신건 아닌데.. 아마도 다른 논문이 더 있지 않을까 싶지만..).
어이쿠야.. 인삼에 관한 것도 있다. 중국애들이 낸 논문인데 2개의 microsatellite loci(CT 12, CA 33) 에 대해 동양의 ginseng과 미국거랑 different allele pattern을 보인단다. 뭐 여튼 난 알고 싶은건 길이니까. 아래 표를 참고하자.
STR이 뭐냐? 위키에 따르면 두개 이상의 nucleotide가 연속적으로 반복되어 있는것. 보통 non-coding intron에서 나타난단다. STRP (STR Polymorphism)은 사람마다 STR loci에 repeat의 수가 다름으로서 생기는 다형성. 법의학적으로 굉장히 많이 사용된단다.
여기까지만 하고 DB부터 검색해 보자.
구글부터 검색을 해보자면.. 여러개 나온다. ATCC, YHRD, ENFSI, STRBase. STRBase가 웹페이지는 허접해 보이나 왠지 좀 academic 해 보이는 관계로 이걸 위주로 보겠다.
STRBase의 DB논문은 2001년에 NAR에 나왔다. 이 DB는 NIST(National Institute of Standards and Technology)에서 관리하는 DB로 뭐 STR DNA marker, 관련 논문 정리, 관련 연구 기관, STR multiplex kits 등 STR 과 관련된 정보는 다 모아놓은 DB라고 할 수 있겠다. 여기까지가 abstract 내용이고
STRBase에서 제공하는 intro를 보자면 forensic DNA typing의 역사서 부터 시작해서 어떤 케이스에 STR이 사용되는지 실험kit, 기계, base call signal, 등등 아주 간략하게 그림만 나와있는데 마지막 슬라이드가 흥미롭다. CODIS(COmbined Dna Index System)라고 FBI의 DB가 있단다. 아래 그림에 나와있는 13 개의 dna loci 에 대해 법의학적으로 DNA profiling을 한다고 한다.
아 근데 이게 좀.. 음 STR이 microsatellite의 하나인 것 같은데.. 위키에서도 보면 두 개의 정의가 merge하도록 suggestion 되어 있고 다음 pdf를 봐도 microsatellite = STS + STR로 되어 있다. 음.. 나와 같은 고민을 한사람이 있었다. 여기 참조뭐.. 여튼
여기까지가 대충 한번 훑어 본거고. 그럼 내가 알아봐야 하는건 무엇인가. 이 STR의 길이의 variation과 maximum length. 음 근데 이거 STRBase에 잘 나와있다. 특히나 cat, dog, cattle의 경우 굉장히 편리하게 table로 정리해놨다. 음.. 사이즈가 대략 maximum 몇백 하는군..
하지만 내가 궁극적으로 찾아야 하는건 뭐냐. 식물에서의 STR. googling 하자마자 나온게 재밌게도 대마초의 STR에 관한 논문. STR이 법의학적으로 사용이 많이 되서 그런지 내가 찾아보면서 본 논문들은 거의 Forensic Science International 에서 나온 듯. abstract를 보자면 93개의 대마를 5개의 STR marker 로 profiling 했단다. 이 5 loci 에서 총 79 개의 allele이 확인됐다. 그담에 AMOVA로 genetic variation을 찾고 fibre crop accession과 drug accession의 variation이 어떻고 뭐하고 뭐한다는데.. 뭐래는 건지 모르겠다(결론은 drug종과 fibre종간의 차이를 확실하게 구분짓지는 못하나 drug 종간에는 사용할만하다 뭐 이런거 같은). 여튼 확인하고자 하는 건 STR의 maximum length. 아래 표가 그것을 의미하는게 아닐까 한다. 얼추 2~300 정도 하는거 같네.
그 다음에 드는 생각이 대마도 있는데 arabidopsis는 없을까? 역시 없을리 없다. 다음 논문을 보자(그닥 최신건 아닌데.. 아마도 다른 논문이 더 있지 않을까 싶지만..).
어이쿠야.. 인삼에 관한 것도 있다. 중국애들이 낸 논문인데 2개의 microsatellite loci(CT 12, CA 33) 에 대해 동양의 ginseng과 미국거랑 different allele pattern을 보인단다. 뭐 여튼 난 알고 싶은건 길이니까. 아래 표를 참고하자.
Thursday, May 19, 2011
genome sequencing 정리
genome assembly를 회사에서 시작할려는지 박사님께서 데이터를 하나 보내왔다(구글에서 찾아진다.. 이 말인즉 누구 박사가 어쩌다 하나 찾아서 참고하라고 보냈나싶다).
사실 예전 부터 genome sequencing 부분에서 조금씩 이해가 안가는 부분이 있었는데 이를 확실하게 정리해본다. 첨 human genome sequencing 했을 때 사용됐던 방법이 shotgun sequencing 이랑 bac-to-bac 방식 거기에 대한 소개는 다음에 잘 나와있다. shotgun이야 익숙한데 bac-to-bac은 익숙하지 않다. 특히나 physical mapping 부분은 이해가 되지 않는다(genetics 수업을 제대로 안들었다). 해서 여기를 참조한다(아 요 physical mapping pdf 어렵다.. ).
처음 데이터에 있는 내용은 음.. 각종 시퀀싱 platform을 섞어서 한거 같은데 Bac-by-Bac도 진행해서 굉장히 사용가능한 정보가 많은거 같은데.. ppt 형식만 봐서는 정확히 이해는 가지 않는다. 논문 method를 보지 않는 이상.
사실 예전 부터 genome sequencing 부분에서 조금씩 이해가 안가는 부분이 있었는데 이를 확실하게 정리해본다. 첨 human genome sequencing 했을 때 사용됐던 방법이 shotgun sequencing 이랑 bac-to-bac 방식 거기에 대한 소개는 다음에 잘 나와있다. shotgun이야 익숙한데 bac-to-bac은 익숙하지 않다. 특히나 physical mapping 부분은 이해가 되지 않는다(genetics 수업을 제대로 안들었다). 해서 여기를 참조한다(아 요 physical mapping pdf 어렵다.. ).
처음 데이터에 있는 내용은 음.. 각종 시퀀싱 platform을 섞어서 한거 같은데 Bac-by-Bac도 진행해서 굉장히 사용가능한 정보가 많은거 같은데.. ppt 형식만 봐서는 정확히 이해는 가지 않는다. 논문 method를 보지 않는 이상.
Tuesday, March 8, 2011
metagenomics
요즘 구제역이다 뭐다 해서 가축들을 죄다 땅에 파 묻는 바람에 지하수에 구제역에 의한 오염이 있지 않나 뭐다나 해서 농진청에서 연구비를 지원하나보다. 덕분에 회사의 가장 힘없는 말단 사원인 난 metagenomics 세계로 뛰어 들게 된다(근데 들어보니 구제역은 바이러스 때문이라는데..). 평소에 metagenomics에 대해 생각이 없었는데.. 예전에 천교수님 발표 할때 들어보고.. 아 꽤 시장이 크구나라고 느낀게 전부인데.. 나는야 까라면 쪼금 반항해보고 결국 까고 마는 말단 사원이다(아.. 연구원이다.. 사원보다 월급 적게 받는).
뭐 덕분에 공부한다고 생각하고 하나하나 정리해보자.
metagenomics 란 무엇인가?
rRNA
metagenome 시퀀싱을 하면 일반적으로 rRNA를 시퀀싱한다 (물론 그냥 gDNA를 culture해서 orf도 prediction하고 protein 시퀀스를 이용해서 functional annotation도 하지만). 아직 까지 내가 아는 지식으론 아마도 그 orgamism들의 구성도를 보기 위해서? 여튼.. 아래 위키 for rRNA explanation
http://en.wikipedia.org/wiki/Ribosomal_RNA ribosomal RNAs는 LSU(large subunit), SSU(small subunit) 으로 구성되어 있는데 prokaryote의 경우 LSU로 50S가 SSU로 30S 가 있고 그 30S를 구성하는 rRNA가 바로 16S rRNA. 보통 rRNA sequencing 중 16S rRNA 시퀀싱을 많이 하는데 그 이유를 생각해 보자면 http://en.wikipedia.org/wiki/16S_ribosomal_RNA 에 마지막에 보면 16S rRNA에 universal primer를 쓸수 있을 정도로 conserved 한 region도 있고 반면에 굉장히 변화가 심한 hypervariable region도 있기 때문에 아마도 species를 구분하기에 적당해서가 아닐까.
참고로 다음 논문도 읽어볼만 할듯하다.
Ribosomal RNA : a key to phylogeny
http://www.fasebj.org/content/7/1/113.full.pdf#page=1&view=FitH
metagenomics 분석 어떻게 해야 하나?
http://mmbr.asm.org/cgi/content/short/72/4/557
metagenomics를 위한 bioinfomatic 가이드라는 제목의 review인데.. 꼭 읽어봐야 할듯. intro 바로 처음에 나오듯이 이 리뷰는 functional metagenomics (특정 activity가 있는 것만 골라내서 cloning 해서 시퀀싱 한거)랑 구분하여 50Mbp 이상의 randomly sampled sequnces를 분석하는 가이드.
관련 데이터 베이스
들어보니 Silva, greengenes, EZ_taxon 이렇게 3개가 가장 많이 사용되는거 같다. 몇개 찾아보니 Silva가 가장 잘 되어 있는 느낌(?)이 드는데 우선 관련 논문
<Silva>
http://nar.oxfordjournals.org/content/35/21/7188.full?keytype=ref&ijkey=pwbw9T96ADMbJBk
일단 이러한 데이터 베이스의 목적은 넘쳐나는 rRNA데이터를 careful inpection, 그러니까 curration을 통해 rRNA가 biodiversity 연구에 도움이 되도록 하는데 있다(unified quality control & alignment of rRNA datasets).
논문을 보니까 rRNA 를 이용한 phylogeny의 연구에 ARB 라는 software와 이를 위한 db를 많이 썼었던걸로 보인다. ARB 말고도 rRNA curation을 위한 3개의 main project를 소개한다(greengenes, RDP,그리고 하나가 european rRNA 데이터 베이스인데 이것이 Silva로 들어 간것으로 보인다.). 아 그리고 하나 greengene에서는 ARB compatible dataset을 가지고 있긴 한데 full length인것만 대해서만. 그리고 요즘은 LSU rRNA도 많이 사용한다네(특히 eukaryote의 경우). intro을 본 결론은 4개의 DB 그중 ARB랑 european rRNA은 Silva로 편입된거 같다.
-Sequence data
Silva의 버젼은 EMBL과 버젼이 똑같다. 곧 RNA와 관련된 키워드 모두 검색해서 EMBL에서 RNA 시퀀스를 가져온다는 말. 그리고 seed alignment를 제공한다는데.. silva의 예전 버젼 격인 ARB에서 release 한것을 그대로 유지한다는데.. 이 seed alignment라는게 뭔지 잘 모르겠다..
-Quality checks
1.unaligned uncleotides 중 300bp보다 짧거나, 2. 2%이상 ambiguities가 있거나, 3. 아니면 2%이상의 homopolymer (homotetramer(homo-4bp)이상)가 있거나, 4. vector랑 5%이상 매치되면 제거. 그리고 이 세가지(2,3,4 항목)의 평균을 구해서 100에서 빼면 이것이 sequence quality가 된다. 이후 필터링 통과한 시퀀스들은 seed alingment에 대해 SINA(silva incremental aligner)에 의해 alignment가 되어 진다. 근데 이 sequence quality로 뭐하는 것인지.. 이미 4가지 항목으로 필터링 했는데 그 뒤에 이 sequence quality를 왜 구하는건지..이 역시 아직 잘 모르겠음
-Aligner
ARB의 suffix tree[1] 방식을 이용해서 seed alignment에서 최대 40개까지 유사한 시퀀스를 찾는다. 이렇게해서 찾아진 시퀀스 들은 partial order graph[2]로 옮기고 이 graph 위에다가 query를 needleman 방식으로 align 한다. 이 과정에서 alignment quality와 basepair score를 구하고 이는 0~100 사이 값으로 normalized 한다. alignment를 마친뒤에 aligned된 bp가 300bp보다 작으면 버린다.
-Anomaly check
이건 seed sequence의 anomaly를 체크 하기 위한 것이거 같은데. pintail이란 프로그램을 사용한단다. seed sequence들 전보를 20 개의 sequence로 된 한 그룹에 대해 pairwise check를 하는데 만약 대부분의 alignment가 anomalous 하게 나오면 seed에서 제거한다는거 같은데.. 저 20개의 sequence가 정확히 뭔지 모르겠다.. 모르는거 투성이네. 에이
-Taxonomy
-Nomenclature
-SSU and LSU rRNA databases for ARB
Ref databases: Parc database의 subset, 1.length cutoff :거의 full length의 시퀀스(최소 1200bp). archaea의 경우 800bp. LSU의 경우 1900bp. 2.alignment curoff : alignment score가 SSU의 경우 50, LSU의 경우 30 이상. 이 뒤에 positional variability filtering이 있는데 잘 이해 안됨.
Parc databases: 위의 quality가 확인된 모든 sequences
[1]suffix tree: 이진트리나 레드블랙트리는 봤어도 suffix tree는 사실 제대로 본적이 없다. 금선생이 추천해준 책에 몇챕터에 걸쳐서 나오는데.. 아.. 역시 모든 지식은 연결된듯하다. 여튼 급한데로 훓어보는데.. http://graphy21.blogspot.com/2011/03/suffix-tree.html
[2]http://bioinformatics.oxfordjournals.org/content/18/3/452.full.pdf#page=1&view=FitH
뭐 덕분에 공부한다고 생각하고 하나하나 정리해보자.
metagenomics 란 무엇인가?
rRNA
metagenome 시퀀싱을 하면 일반적으로 rRNA를 시퀀싱한다 (물론 그냥 gDNA를 culture해서 orf도 prediction하고 protein 시퀀스를 이용해서 functional annotation도 하지만). 아직 까지 내가 아는 지식으론 아마도 그 orgamism들의 구성도를 보기 위해서? 여튼.. 아래 위키 for rRNA explanation
http://en.wikipedia.org/wiki/Ribosomal_RNA ribosomal RNAs는 LSU(large subunit), SSU(small subunit) 으로 구성되어 있는데 prokaryote의 경우 LSU로 50S가 SSU로 30S 가 있고 그 30S를 구성하는 rRNA가 바로 16S rRNA. 보통 rRNA sequencing 중 16S rRNA 시퀀싱을 많이 하는데 그 이유를 생각해 보자면 http://en.wikipedia.org/wiki/16S_ribosomal_RNA 에 마지막에 보면 16S rRNA에 universal primer를 쓸수 있을 정도로 conserved 한 region도 있고 반면에 굉장히 변화가 심한 hypervariable region도 있기 때문에 아마도 species를 구분하기에 적당해서가 아닐까.
참고로 다음 논문도 읽어볼만 할듯하다.
Ribosomal RNA : a key to phylogeny
http://www.fasebj.org/content/7/1/113.full.pdf#page=1&view=FitH
metagenomics 분석 어떻게 해야 하나?
http://mmbr.asm.org/cgi/content/short/72/4/557
metagenomics를 위한 bioinfomatic 가이드라는 제목의 review인데.. 꼭 읽어봐야 할듯. intro 바로 처음에 나오듯이 이 리뷰는 functional metagenomics (특정 activity가 있는 것만 골라내서 cloning 해서 시퀀싱 한거)랑 구분하여 50Mbp 이상의 randomly sampled sequnces를 분석하는 가이드.
관련 데이터 베이스
들어보니 Silva, greengenes, EZ_taxon 이렇게 3개가 가장 많이 사용되는거 같다. 몇개 찾아보니 Silva가 가장 잘 되어 있는 느낌(?)이 드는데 우선 관련 논문
<Silva>
http://nar.oxfordjournals.org/content/35/21/7188.full?keytype=ref&ijkey=pwbw9T96ADMbJBk
일단 이러한 데이터 베이스의 목적은 넘쳐나는 rRNA데이터를 careful inpection, 그러니까 curration을 통해 rRNA가 biodiversity 연구에 도움이 되도록 하는데 있다(unified quality control & alignment of rRNA datasets).
논문을 보니까 rRNA 를 이용한 phylogeny의 연구에 ARB 라는 software와 이를 위한 db를 많이 썼었던걸로 보인다. ARB 말고도 rRNA curation을 위한 3개의 main project를 소개한다(greengenes, RDP,그리고 하나가 european rRNA 데이터 베이스인데 이것이 Silva로 들어 간것으로 보인다.). 아 그리고 하나 greengene에서는 ARB compatible dataset을 가지고 있긴 한데 full length인것만 대해서만. 그리고 요즘은 LSU rRNA도 많이 사용한다네(특히 eukaryote의 경우). intro을 본 결론은 4개의 DB 그중 ARB랑 european rRNA은 Silva로 편입된거 같다.
-Sequence data
Silva의 버젼은 EMBL과 버젼이 똑같다. 곧 RNA와 관련된 키워드 모두 검색해서 EMBL에서 RNA 시퀀스를 가져온다는 말. 그리고 seed alignment를 제공한다는데.. silva의 예전 버젼 격인 ARB에서 release 한것을 그대로 유지한다는데.. 이 seed alignment라는게 뭔지 잘 모르겠다..
-Quality checks
1.unaligned uncleotides 중 300bp보다 짧거나, 2. 2%이상 ambiguities가 있거나, 3. 아니면 2%이상의 homopolymer (homotetramer(homo-4bp)이상)가 있거나, 4. vector랑 5%이상 매치되면 제거. 그리고 이 세가지(2,3,4 항목)의 평균을 구해서 100에서 빼면 이것이 sequence quality가 된다. 이후 필터링 통과한 시퀀스들은 seed alingment에 대해 SINA(silva incremental aligner)에 의해 alignment가 되어 진다. 근데 이 sequence quality로 뭐하는 것인지.. 이미 4가지 항목으로 필터링 했는데 그 뒤에 이 sequence quality를 왜 구하는건지..이 역시 아직 잘 모르겠음
-Aligner
ARB의 suffix tree[1] 방식을 이용해서 seed alignment에서 최대 40개까지 유사한 시퀀스를 찾는다. 이렇게해서 찾아진 시퀀스 들은 partial order graph[2]로 옮기고 이 graph 위에다가 query를 needleman 방식으로 align 한다. 이 과정에서 alignment quality와 basepair score를 구하고 이는 0~100 사이 값으로 normalized 한다. alignment를 마친뒤에 aligned된 bp가 300bp보다 작으면 버린다.
-Anomaly check
이건 seed sequence의 anomaly를 체크 하기 위한 것이거 같은데. pintail이란 프로그램을 사용한단다. seed sequence들 전보를 20 개의 sequence로 된 한 그룹에 대해 pairwise check를 하는데 만약 대부분의 alignment가 anomalous 하게 나오면 seed에서 제거한다는거 같은데.. 저 20개의 sequence가 정확히 뭔지 모르겠다.. 모르는거 투성이네. 에이
-Taxonomy
-Nomenclature
-SSU and LSU rRNA databases for ARB
Ref databases: Parc database의 subset, 1.length cutoff :거의 full length의 시퀀스(최소 1200bp). archaea의 경우 800bp. LSU의 경우 1900bp. 2.alignment curoff : alignment score가 SSU의 경우 50, LSU의 경우 30 이상. 이 뒤에 positional variability filtering이 있는데 잘 이해 안됨.
Parc databases: 위의 quality가 확인된 모든 sequences
[1]suffix tree: 이진트리나 레드블랙트리는 봤어도 suffix tree는 사실 제대로 본적이 없다. 금선생이 추천해준 책에 몇챕터에 걸쳐서 나오는데.. 아.. 역시 모든 지식은 연결된듯하다. 여튼 급한데로 훓어보는데.. http://graphy21.blogspot.com/2011/03/suffix-tree.html
[2]http://bioinformatics.oxfordjournals.org/content/18/3/452.full.pdf#page=1&view=FitH
Thursday, March 3, 2011
TSS, TxStart, CDSstart,
맨날 까먹고 헷갈리고.. 에이..
TSS : transcription start site. 그러니까 5'UTR 부분부터.
http://en.wikipedia.org/wiki/Transcription_start_site
txStart : TSS 랑 같은말
cdsStart : coding region start site로 start codon 부터 시작. 곧 protein 시작 부분.
그런데 헷갈렸던 것이 위 그림. 어떤 exon은 wholly or partially 5' UTR 이기때문에 exon start site가 cds start 랑 같지 않다.
TSS : transcription start site. 그러니까 5'UTR 부분부터.
http://en.wikipedia.org/wiki/Transcription_start_site
txStart : TSS 랑 같은말
cdsStart : coding region start site로 start codon 부터 시작. 곧 protein 시작 부분.
그런데 헷갈렸던 것이 위 그림. 어떤 exon은 wholly or partially 5' UTR 이기때문에 exon start site가 cds start 랑 같지 않다.
Thursday, January 20, 2011
1000 genome supplementary note
nature에 나온 1000 genome 논문의 supplementary information을 정리한다.
2. Samples
YRI (Yoruba in Ibadan, Nigeria), CEU (ancestry from Northern and Western Europe), CHB (Han Chinese in Beijing, Chian), JPT (Japanese in Tokyo, Japan), LWK (from the Luhya in Webuye, Kenya), TSI (Toscani in Italia), CHD(the Chinese in Metropolitan Denver, CO, USA)
4. Read mapping and generation of BAM files
-quality recalculated -> remap -> merge lanes from the same library (Picard MergeSamFiles) -> remove duplicate (samtools : rmdup for paired end, rmdupse for single end) -> merge libraries to the plaform level -> remove duplicate (Picard MarkDuplicates)
4.1 Reference genome
-NCBI36, revised Cambridge reference sequence instead of mtDNA. sex-specific reference (Y chr only for male, psudoautosomal region masked in Y chr)
4.2 Mapping of Illumina Data
-Maq v0.7 -u -a 1000
4.5 Recalibration of Base Quality Values
-recalibrate qulity after initial alignment. this algorithm(covariate-aware base quality recalibration algorithm) is implemented in GATK software.
http://www.broadinstitute.org/gsa/wiki/index.php/Base_quality_score_recalibration
-effect of recalibaration : the total number of variants called decreased by 2.8%. changing Ti/Tv ratio from 1.07 to 1.96 (true variants around 2. random 0.5)
Ti/Tv ratio :
http://www.cbs.dtu.dk/staff/dave/roanoke/genetics980415f.htm
http://paup.csit.fsu.edu/paupfaq/paupans.html
4.6 Comparison of Read Data to known HapMap Genotypes
-genotype log likelihood (samtools pileup -g) was used for matching expected genotype, and if the best genotype did not seperate well from the others(1.2 separation), then removed.
likelihood:
http://www.aistudy.co.kr/math/likelihood.htm
genotype likelihood : maq paper(http://graphy21.blogspot.com/2011/01/maq.html)
5.SNP calling
-maq에서 나온 genotype likelihood(GLij(g) = P(Bij,Qij | Gij =g))를 이용해서 snp를 call한다.
P(Gij = g|Bij,Qij) = P(Bij,Qij | Gij = g) P(Gij=g) / Kij , Kij = Σg P(Bij, Qij| Gij = g) P(Gij = g)
말로 풀어서 다시 말하면 maq에서 나온 공식으로 genotype likelihood를 구하고 bayesian공식으로 poterior probability, 즉 데이터가 나왔을때 어떤 genotyp이냐를 추즉한다.
이렇게 snp가 call되면 post-processing step으로 false positive를 제거하고 VCF (variant call format) 형식으로 저장한다.
-post-processing filtering {
--expected depth보다 너무 낮거나 높은거(평균 depth의 반 or 두배), 아마도 CNV에 의한 paralog로 잘못 mapping된거라 생각되서
--snp call 부위의 local realignment, indel에 의한 misalignment를 방지 하기 위해(보통 gap open penalty가 mismatch보다 크다)
--poor mapping quality 제거 , reference 자체가 완벽하지 않기때문에 unrepresented region에서 나온 read가 잘못 맵핑될수 있다(경험상 잘못 mapping되는 region을 6.1에 있다). }
5.1 Low-Coverage SNP calling
5.3 Exon project SNP calls
Thursday, November 18, 2010
phylogeny tree
영건씨와의 프로젝트에서 내가 맡은 부분인 distance matrix에서 phylogeny tree를 그리는 것을 수행하기 위해.. 일주일 정도 나름 고생했다.. 아..
나름 고생해서 얻은 것들은 대충 적는다.
phylogeny tree를 그리는 방법에는 UPGMA (biological sequence analysis(BSA) 책에 보면 이름만 거창하다는 식으로 intro를 시작한다), neighbor joining, parsimony 가 있는데 이번에 선택한 방법은 neighbor joining (NJ). 왜냐고 물어보면... 음.. FFP를 사용한 reference가 되는 sims의 논문이 FFP를 구하고 나서 이를 NJ 방법으로 tree를 그렸기 때문에?(좀 구차한거 같은 느낌이.. 사실 위의 방법들의 장단을 봐야 하는데)..
NJ 방법이 알고보니 unrooted tree를 그리는 것이다 (아.. 것도 모르고 마지막 두개 남은 node에서 error가 나는걸보고 알고리즘이 이상하다고 느꼇음). BSA 책에 보면 rooted로 만드는 2가지 방법을 소개하는데, 하나는 완전 outgroup인 것을 넣어서 그걸 root로 삼으라는 것과 연결선의 가장 긴 chain의 midpoint를 root로 하라는 것인데.. outgroup을 억지로 넣어주면 edge의 값이 조금씩 변하는게 맘에 들지 않고 해서 그냥.. 나름 unroot인데 root인 마냥 나오게끔..
그리고 결과 파일을 newick format으로 만드는 작업은 추가 해야 겠다. 다른 프로그램 (phylip 같은) 것에서도 사용할수 있게.
아래는 이번 작업하면서 무자비하게 search 한것 모음
---------------------------------------------------------
정리해 보자면 현재 나에게 떨어진 해당 과제가.. FFP를 이용한 시퀀스 사이의 similarity matrix가 있을때 이를 phylogeny tree로 그려야 하는것.
이는 clustalw가 모든 pairwise 시퀀스를 다이나믹 프로그래밍으로 alignment 해서 similarity를 구한다는 것 이외에는 그 뒤 과정이 내가 해야 하는 것과 동일하므로. clustalw를 참조하기로 한다.
우선 clustalw가 neighbor joining 방식으로 similarity 트리를 그리므로 neighbor joining 한번 점검하고... 역시 clustalw 처럼 dnd 형식을 최종 output으로 만들고 싶기 때문에 neighbor joining 이후에 dnd 파일 만들기를 check 해야 할것이다.
기본적으로 multiple sequence alignment와 phylogeny tree 그리기가 무엇이지를 알기위해서는 http://www.cbs.dtu.dk/courses/humanbio/2010/exercises/ExMulPhyl/Ex_Phylo.php에 나오는 HIV의 phylogeny tree 그리기 예제를 살펴 봐야 전체적인 그림이 나올것 같고..
그런데 dnd 파일이 꼭 phylogeny tree를 나타내는 것이 아니라는 내용이 http://www.ebi.ac.uk/Tools/clustalw2/faq.html에 guide tree와 phylogeny tree의 차이가 무엇이냐라는 것에 있고.. 좀 헷갈리는데.. 알아봐야 할거 같네.. 그냥 clustalw2를 돌려서 나오는 dnd가 guide tree인지 아니면 phylogeny tree인지(아무래도 이건 최종으로 나오는 거라 phylogeny tree 인거 같지만)
음 보아하니.. 순수하게 pairwise alignment 한거가지고 나온건 guide tree인것이고 이거로 multiple alignment 해서 나온 결과로 그린것이 phylogeny tree인 것인데.. 기본적으로 clustalw의 manual을 훑어볼필요가 있는듯
-HMMER3는 multiple alignment file format을 clustalw2에서 나온 aln을 인식하지 못한다. 그래서 biopython의 AlignIO.convert 를 이용해서 HMMER3가 인식할수 있는 stockholm 파일로 변환해서 사용한다(AlignIO.convert('temp.aln','clustal','temp.sto','stockholm')).
-clustalw2는 phylip의 file format으로 output file 생성 가능
clustalw의 과정
step1 : 가능한 모든 pair의 시퀀스를 align 하고 distance(mismatch position/non-gapped position)를 정한다. 그 다음 distance matrix를 만든다.
step2 : neighbor joining method를 이용해서 similarity tree를 그린다.
step3 : 위의 similarity tree를 참조해서 가까운 시퀀스부터 하나씩 combine 해서 multiple alignment 를 하는데.. 위의 similarity tree의 root 로부터의 거리를 weight를 갖게 된다. 그러니 같은 banch 안에 들어가는 시퀀스는 그 weight를 나눠 갖게 된다.
-check list
1.neighbor joining method
2. dnd file format (newick)
Monday, November 1, 2010
각종 non-coding RNA
요즘 논문을 보면 각종 RNA에 대한 논문들이 많은데.. 그래서 이러한 각종 RNA에 대해 정리를 해야 한다는 생각이 들어 단순히 wiki를 링크 걸어본다. 읽어보고 정리는 언제 할지 미지수지만.. 시작이 반이라고..
non-coding RNA : http://en.wikipedia.org/wiki/Non-coding_RNA
miRNA : http://en.wikipedia.org/wiki/MicroRNA
piRNA : http://en.wikipedia.org/wiki/Piwi-interacting_RNA
snoRNA : http://en.wikipedia.org/wiki/SnoRNA
snRNA: http://en.wikipedia.org/wiki/Small_nuclear_RNA
lincRNA : http://en.wikipedia.org/wiki/Long_non-coding_RNA
lincRNA는 large intergenic non coding RNA 의 준말인데.. 음 long non coding RNA의 한 종류이겠지...
non-coding RNA : http://en.wikipedia.org/wiki/Non-coding_RNA
miRNA : http://en.wikipedia.org/wiki/MicroRNA
piRNA : http://en.wikipedia.org/wiki/Piwi-interacting_RNA
snoRNA : http://en.wikipedia.org/wiki/SnoRNA
snRNA: http://en.wikipedia.org/wiki/Small_nuclear_RNA
lincRNA : http://en.wikipedia.org/wiki/Long_non-coding_RNA
lincRNA는 large intergenic non coding RNA 의 준말인데.. 음 long non coding RNA의 한 종류이겠지...
Monday, August 23, 2010
homologous recombination
우선 이번 포스팅은 코리안으로 하겠다.
오늘 science 잡지에 "re-replication may be a contributor to gene copy number changes" 라는 제목으로 논문이 실렸다. 그 메커니즘은 오른쪽과 같다. NAHR(non-allelic homologous recombination)에 의해 gene copy number 가 달라진다는 내용인듯하다 (그림만 보고 읽진 않았다). 이 그림을 보고 NAHR이 무엇인가를 찾아보게 되었다. 다름 아니라 하나의 allele에서 일어나는 HR(homologous recombination).
그렇다면 HR은 무엇인가? recombination은 재조합으로 예전 생물학 시간에 들었던 것이 얼핏 기억이 난다. 우선 두군데에서 정보를 찾았다.
sanger 와 wikipedia.
생거에서 말하는 정보는 매우 적고 좀더 자세히나온 위키 피디아를 본다.핵심 적인 내용(왼쪽 그림)은 double strand break repairs 과정의 하나의 pathway인 DSBR pathway에 의해 HR이 생기고 그 결과 crossover 내지는 gene conversion이 생긴다는 것이다.
double Holliday junctions이 nicking endonuclease에 의해 horizontal resolution(sanger site에서의 표현을 빌리자면) 에 의하면 gene conversion이 일어나고 vertical resolution에 의해 cross-over가 일어난다.
문제는 일반적으로 gene conversion을 찾아보면 그림이 오른쪽과 같은데 위의 과정에는 gene conversion이 일어나면 한쪽 duplex에는 다른 한쪽의 duplex에서 온 dna 조각이 double strand로 통으로 들어가는게 맞지만 gene conversion이 일어나지 않은 duplex에도 one strand로 다른 쪽의 dna 조각이 들어가게 되는데 오른쪽의 그림에서는 한쪽은 전혀 dna 가닥이 섞이지 않게 표현되어 있다. 이는 내가 잘못 이해한것인가 아니면 편의상 그림을 오른쪽과 같이 그린것인가?
anyone can answer this problem????
오늘 science 잡지에 "re-replication may be a contributor to gene copy number changes" 라는 제목으로 논문이 실렸다. 그 메커니즘은 오른쪽과 같다. NAHR(non-allelic homologous recombination)에 의해 gene copy number 가 달라진다는 내용인듯하다 (그림만 보고 읽진 않았다). 이 그림을 보고 NAHR이 무엇인가를 찾아보게 되었다. 다름 아니라 하나의 allele에서 일어나는 HR(homologous recombination).
그렇다면 HR은 무엇인가? recombination은 재조합으로 예전 생물학 시간에 들었던 것이 얼핏 기억이 난다. 우선 두군데에서 정보를 찾았다.
sanger 와 wikipedia.
생거에서 말하는 정보는 매우 적고 좀더 자세히나온 위키 피디아를 본다.핵심 적인 내용(왼쪽 그림)은 double strand break repairs 과정의 하나의 pathway인 DSBR pathway에 의해 HR이 생기고 그 결과 crossover 내지는 gene conversion이 생긴다는 것이다.
double Holliday junctions이 nicking endonuclease에 의해 horizontal resolution(sanger site에서의 표현을 빌리자면) 에 의하면 gene conversion이 일어나고 vertical resolution에 의해 cross-over가 일어난다.
문제는 일반적으로 gene conversion을 찾아보면 그림이 오른쪽과 같은데 위의 과정에는 gene conversion이 일어나면 한쪽 duplex에는 다른 한쪽의 duplex에서 온 dna 조각이 double strand로 통으로 들어가는게 맞지만 gene conversion이 일어나지 않은 duplex에도 one strand로 다른 쪽의 dna 조각이 들어가게 되는데 오른쪽의 그림에서는 한쪽은 전혀 dna 가닥이 섞이지 않게 표현되어 있다. 이는 내가 잘못 이해한것인가 아니면 편의상 그림을 오른쪽과 같이 그린것인가?
anyone can answer this problem????
darwin's evolution theory
Several months ago, I was fascinated with evolution theory. So I found many documents and web sites for a week and tried to read most of them. But because whether my limitation on interpreting or lack of detailed explanation, my understanding to evolution theory is not clear at that time.
I found this book on last week by chance and read it on last weekend. Through this book I can do arrange my thinking about evolution theory and its' history.
Especially a section which was written by Motoo Kimura thirty years ago who make neutral theory of molecular evolution was really impressive. I read that part as if I heard lecture from Kimura.
I believe that ALL analysis of biological data is based on evolution theory and the biologists who don't have this concept are scientists without spirit.
As mentioned in Darwin 2.0 (the book? which is written by professor in Ewha univ.), Darwin's theory is the one which can be considered as the greatest theory in biology and it's greatness can be comparable with theory of relativity from physics.
I can found theory about evolution and meaning of population genetics and application of evolution. Of course although these are not that concrete, it's enough to set the basic concept of these things.
I found this book on last week by chance and read it on last weekend. Through this book I can do arrange my thinking about evolution theory and its' history.
Especially a section which was written by Motoo Kimura thirty years ago who make neutral theory of molecular evolution was really impressive. I read that part as if I heard lecture from Kimura.
I believe that ALL analysis of biological data is based on evolution theory and the biologists who don't have this concept are scientists without spirit.
As mentioned in Darwin 2.0 (the book? which is written by professor in Ewha univ.), Darwin's theory is the one which can be considered as the greatest theory in biology and it's greatness can be comparable with theory of relativity from physics.
I can found theory about evolution and meaning of population genetics and application of evolution. Of course although these are not that concrete, it's enough to set the basic concept of these things.
Saturday, July 17, 2010
pseudo gene
I felt the need to clarify about pseudo gene while I organized an annotation of zymo. Therefore in this post I will review about pseudo gene, what it is and how it was happened.
All information comes from http://en.wikipedia.org/wiki/Pseudogene
Pseudogene : non-expressed or defunctional relatives of known gene (homology to a known gene, nonfunctionality).
1.homology to a known gene : usually sequence identity is between 40% and 100%.
2.nonfunctionality : if any one of the common processes, which are needed in making functional product from DNA, such as transcription and pre-mRNA processing fails, this will be nonfunctional. In high-throughput psudogene identification, modification of stop codon and frameshifts are the most common factors.
Types and origin of pseudogenes
1. Processed (retrptransposed) pseudogenes : retrotransposition event is common in mammals and in case of human 30%~44% of genome is composed of this. In retrotransposition, mRNA transcript of a gene is reverse transcribed back into DNA become pseudogene. This kind of pseudogene have features of cDNA like poly-A tail, introns spliced out and lack of promoter.
2. Non-processed pseudogenes : Through gene duplication A copy of gene is made and subsequent mutation make them as pseudogene. This kind of pseudogene have the most feature of genes.
3. Disable genes (unitary pseudogene) : same mechanism (i.e. mutation) which make non-processed pseudogene happened to genes without duplication and became fixed in population.
All information comes from http://en.wikipedia.org/wiki/Pseudogene
Pseudogene : non-expressed or defunctional relatives of known gene (homology to a known gene, nonfunctionality).
1.homology to a known gene : usually sequence identity is between 40% and 100%.
2.nonfunctionality : if any one of the common processes, which are needed in making functional product from DNA, such as transcription and pre-mRNA processing fails, this will be nonfunctional. In high-throughput psudogene identification, modification of stop codon and frameshifts are the most common factors.
Types and origin of pseudogenes
1. Processed (retrptransposed) pseudogenes : retrotransposition event is common in mammals and in case of human 30%~44% of genome is composed of this. In retrotransposition, mRNA transcript of a gene is reverse transcribed back into DNA become pseudogene. This kind of pseudogene have features of cDNA like poly-A tail, introns spliced out and lack of promoter.
2. Non-processed pseudogenes : Through gene duplication A copy of gene is made and subsequent mutation make them as pseudogene. This kind of pseudogene have the most feature of genes.
3. Disable genes (unitary pseudogene) : same mechanism (i.e. mutation) which make non-processed pseudogene happened to genes without duplication and became fixed in population.
Wednesday, July 14, 2010
biological classification
While I handle NGS data of microbe like Zymomonas and Paenilbacillus, because of name of microbes, I got a thought to study classification of livings. As you know, like our name, they also have last name and middle and first (maybe this is stupid comparison). I thought that if I can know which last name is in long name of microbe and where they belong through that name, that will be very helpful to me to understand more concretely or feel more intimately with them. Being familiar is really important to do something with, right? So start!
- Biological classification -
Biological classification is a form of scientific taxonomy(<->folk taxonomy ; bugs, ducks). 7 ranks are defined by international nomenclature codes.
Carolus Linnaeus's work is a being base of modern classification. What I understand in wikipedia (http://en.wikipedia.org/wiki/Biological_classification) is only that he shorten the name of living.
Since the 1960s, based on Darwinian principle which believe every species have come from common ancestor, cladistic taxonomy has emerged. This arrange taxon based on phylogenetic tree. This will be posted.
So , I found another webpage for bacterial taxonomy and I think this will be more helpful in practical to me (http://www.microbiologybytes.com/iandi/3a.html).
- Bacteria-Basic Facts -
What are bacteria? lack of organelles, no nucleus, circular DNA chromosome, maybe cell surface is the most complex region.
Gram stain appearances of medically important bacteria
Gram strain (purple) have been used to dye bacteria for identification by microscope. If bacteria only have cell wall made by peptideglycan, it keep the purple color(gram positive). But the bacteria which have extra cell membrane composed of phospholipid lose the purple color(gram negative). In that case, red dye is used for dyeing.
This site is focused on bacteria of medical importance and the information for naming of bacteria isn't sufficient.
Therefore, I found the other web-site (http://www.bacterio.cict.fr/classification.html). I think this will be last destination.
- Classification, taxonomy and systematics of prokaryotes (bacteria) -
- Additional information -
Strain, clone and species : comments on three basic concepts of bacteriology (http://jmm.sgmjournals.org/cgi/reprint/49/5/397.pdf)
- Biological classification -
Biological classification is a form of scientific taxonomy(<->folk taxonomy ; bugs, ducks). 7 ranks are defined by international nomenclature codes.
Carolus Linnaeus's work is a being base of modern classification. What I understand in wikipedia (http://en.wikipedia.org/wiki/Biological_classification) is only that he shorten the name of living.
Since the 1960s, based on Darwinian principle which believe every species have come from common ancestor, cladistic taxonomy has emerged. This arrange taxon based on phylogenetic tree. This will be posted.
So , I found another webpage for bacterial taxonomy and I think this will be more helpful in practical to me (http://www.microbiologybytes.com/iandi/3a.html).
- Bacteria-Basic Facts -
What are bacteria? lack of organelles, no nucleus, circular DNA chromosome, maybe cell surface is the most complex region.Gram stain appearances of medically important bacteria
Gram strain (purple) have been used to dye bacteria for identification by microscope. If bacteria only have cell wall made by peptideglycan, it keep the purple color(gram positive). But the bacteria which have extra cell membrane composed of phospholipid lose the purple color(gram negative). In that case, red dye is used for dyeing.
This site is focused on bacteria of medical importance and the information for naming of bacteria isn't sufficient.Therefore, I found the other web-site (http://www.bacterio.cict.fr/classification.html). I think this will be last destination.
- Classification, taxonomy and systematics of prokaryotes (bacteria) -
- Additional information -
Strain, clone and species : comments on three basic concepts of bacteriology (http://jmm.sgmjournals.org/cgi/reprint/49/5/397.pdf)
Subscribe to:
Posts (Atom)









