NCBI submission

From Cbcb
Revision as of 18:33, 17 August 2007 by Dpuiu (talk | contribs) (→‎Procedure)
Jump to navigation Jump to search

NCBI

BankIt

seqin: standalone application

WGS

AA

 Compressed archive containing 2 files: ASSEMBLY.xml , MD5 
 Accepted only by uploading to NCBI FTP server.
   server: ftp-private.ncbi.nlm.nih.gov
   login: umd_trc
   passwd: 
   center: UMD   
   description: University of Maryland
 ASSEMBLY XML Schema png 
 ASSEMBLY XML Schema xsd 

TA

Procedure

Files:

.contig, .seq, .qual, .ti2seq_name

Steps:

1. makeConinfo ASSEMBLY.coninfo
 $ more ASSEMBLY.coninfo
 <coninfo>
 <meta name='center'>UMD</meta>
 <meta name='db'>Xoo</meta>
 <meta name='desc'>Xanthomonas oryzae pv. oryzae strain PXO99A</meta>
 <meta name='object'>ASSEMBLY</meta>
 <meta name='species_code'>Xanthomonas oryzae pv. oryzae strain PXO99A</meta>
 <meta name='structure'>Chromosome</meta>
 <meta name='subtype'>NEW</meta>
 <meta name='taxid'>360094</meta>
 <contig id="1106158952778_stitched" conformation="CIRCULAR" subtype="NEW"/>
 <file src="Xoo.contig"/>
 <seq src="Xoo.seq"/>
 <qual src="Xoo.qual"/>
 <idmap  src="Xoo.ti2seq_name" direction="FORWARD"/>
 </coninfo>
2. buildAssemblyArchive ASSEMBLY.coninfo --prompt --subname umd-20070816-125223
 problems:
    * submitter_reference="tigr...." : replace tigr with umd
    * conformation: always LINEAR    : replace LINEAR with CIRCULAR
    * taxid: not recognized          : replace <taxid>id</taxid> with <organism descriptor="TAXID">id</organism>
 $ ls -1 umd-20070816-125223*
 umd-20070816-125223.tar.gz
 umd-20070816-125223/
  1106158952778_stitched_20070817-141849.con
  1106158952778_stitched_20070817-141849.congap
  ASSEMBLY.xml
  MANIFEST
3. validate:
 oXygen: software used by NCBI; license required
 xmllint: open source
 $ xmllint --schema ASSEMBLY.xsd umd-20070816-125223/ASSEMBLY.xml > /dev/null
 umd-20070816-125223/ASSEMBLY.xml validates