Macaque assembly and gene annotation

Assembly

Mmul_1 is a preliminary assembly of the Indian-origin rhesus monkey, Macaca mulatta using whole genome shotgun (WGS) reads from small and medium insert clones. Several WGS libraries, with inserts of 2-4 kb and 10 kb, fosmids with ~35kb inserts, and BACs with 180kb inserts were used to produce the data.

The release was produced by the Macaque Genome Sequencing Consortium, led by the Baylor College of Human Medicine, melding three separate complementary assemblies (created using the Atlas, Celera and PCAP systems). This involved iteratively splitting likely chimeric scaffolds and joining together existing scaffolds where possible. Chimeric scaffolds (<100 total) were identified by breaks in synteny with the human genome, which were confirmed to be artefacts by the other assemblies.

This is a draft sequence and may contain errors so users should exercise caution. Typical errors in draft genome sequences include misassemblies of repeated sequences, collapses of repeated regions, and unmerged overlaps (e.g. due to polymorphisms) creating artificial duplications. However base accuracy in contigs (contiguous blocks of sequence) is usually very high with most errors near the ends of contigs. [More about the assembly].

Gene annotation

The gene set for macaque was built using the Ensembl pipeline. The species-specific resources for macaque are relatively limited, so we decided to take a combined approach utilizing macaque's great similarity to human to aid our annotation efforts. The gene structures are mainly based on alignments to human and macaque protein data. Both macaque and human cDNAs were used to add UTR structures, and finally gene predictions based on Uniprot proteins and human cdnas were used to fill gaps in the annotation.

More information

General information about this species can be found in Wikipedia.

Statistics

Summary

AssemblyMMUL 1.0, Feb 2006
Database version81.10
Base Pairs3,093,871,206
Golden Path Length3,097,179,960
Genebuild byEnsembl
Genebuild methodFull genebuild
Genebuild startedJan 2006
Genebuild releasedAug 2006
Genebuild last updated/patchedMay 2010

Gene counts

Coding genes21,905
Non coding genes6,579
Small non coding genes5,455
Misc non coding genes1,124
Pseudogenes1,762
Gene transcripts44,725

Other

Genscan gene predictions125,893
Short Variants3,123,522
Structural variants123