<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">APM</journal-id><journal-title-group><journal-title>Advances in Pure Mathematics</journal-title></journal-title-group><issn pub-type="epub">2160-0368</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/apm.2013.37A002</article-id><article-id pub-id-type="publisher-id">APM-38482</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Variant Map System to Simulate Complex Properties of DNA Interactions Using Binary Sequences
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>effrey</surname><given-names>Zheng</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Weiqiong</surname><given-names>Zhang</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jin</surname><given-names>Luo</given-names></name><xref ref-type="aff" rid="aff3"><sup>3</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Wei</surname><given-names>Zhou</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ruoyu</surname><given-names>Shen</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Software, Yunnan University, Kunming, China</addr-line></aff><aff id="aff3"><addr-line>School of Life Sciences, Yunnan University, Kunming, China</addr-line></aff><aff id="aff2"><addr-line>School of Software and Microelectronics, Peking University, Beijing, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>conjugatesys@gmail.com(EZ)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>23</day><month>10</month><year>2013</year></pub-date><volume>03</volume><issue>07</issue><fpage>5</fpage><lpage>24</lpage><history><date date-type="received"><day>August</day>	<month>7,</month>	<year>2013</year></date><date date-type="rev-recd"><day>September</day>	<month>11,</month>	<year>2013</year>	</date><date date-type="accepted"><day>September</day>	<month>28,</month>	<year>2013</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   Stream cipher, DNA cryptography and DNA analysis are the most important R&amp;D fields in both Cryptography and Bioinformatics. HC-256 is an emerged scheme as the new generation of stream ciphers for advanced network security. From a random sequencing viewpoint, both sequences of HC-256 and real DNA data may have intrinsic pseudo-random properties respectively. In a recent decade, many DNA sequencing projects are developed on cells, plants and animals over the world into huge DNA databases. Researchers notice that mammalian genomes encode thousands of large noncoding RNAs (lncRNAs), interact with chromatin regulatory complexes, and are thought to play a role in localizing these complexes to target loci across the genome. It is a challenge target using higher dimensional visualization tools to organize various complex interactive properties as visual maps. The Variant Map System (VMS) as an emerging scheme is systematically proposed in this paper to apply multiple maps that used four Meta symbols as same as DNA or RNA representations. System architecture of key components and core mechanism on the VMS are described. Key modules, equations and their I/O parameters are discussed. Applying the VM System, two sets of real DNA sequences from both sample human (noncoding DNA) and corn (coding DNA) genomes are collected in comparison with pseudo DNA sequences generated by HC-256 to show their intrinsic properties in higher levels of similar relationships among relevant DNA sequences on 2D maps. Sample 2D maps are listed and their characteristics are illustrated under controllable environment. Visual results are briefly analyzed to explore their intrinsic properties on selected genome sequences. 
 
</p></abstract><kwd-group><kwd>Pseudo-Random Number Generator; Stream Cipher; HC-256; Binary to DNA; Pseudo DNA Sequence; Large Noncoding; DNA Analysis; 2D Map; Visual Distribution; Variant Map System</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Stream ciphers [1,2] play a key role in modern network security [3,4] especially in multimedia network environments; its core component—pseudo random number generation mechanism [5-7]—takes the central position in modern cryptography [8,9]. Associated with advanced development of bioinformatics, advanced DNA sequencing and analyzing techniques [10,11] have significantly progressed over the past decade.</p><sec id="s1_1"><title>1.1. DNA Cryptography</title><p>DNA cryptography makes joined research in the field of DNA computing and cryptography. Scholars over the world focused on this field and different results are published such as simulating DNA evolution [<xref ref-type="bibr" rid="scirp.38482-ref12">12</xref>], DNA pseudorandom number generator [13-16], DNA cryptography [9,17,18] and so on. However in current situation, DNA cryptography is still at an earlier stage as an emerging area of advanced cryptography.</p><p>In typical results of DNA cryptography on encrypttion, different coding schemes could be randomly selected. E.g. the algorithm in paper [<xref ref-type="bibr" rid="scirp.38482-ref17">17</xref>] applies an encoding formula to express the plaintext on DNA sequence: {00→C, 01→T, 10→A, 11→G}; however in paper [<xref ref-type="bibr" rid="scirp.38482-ref18">18</xref>], the same author uses the coding formula {00→A, 01→T, 10→C, 11→G} for the plaintext on DNA sequence. In encryption environment, all 4! = 24 possible encoding methods could be equally used in different applications.</p></sec><sec id="s1_2"><title>1.2. Stream Cipher HC-256</title><p>Stream ciphers are an important class of encryption algorithms. A stream cipher is a symmetric cipher which operates with a time-varying transformation on individual plaintext digits. The ECRYPT Stream Cipher Project (eSTREAM) [<xref ref-type="bibr" rid="scirp.38482-ref1">1</xref>] was a multi-year effort, running from 2004-2008, to promote the design of efficient and compact stream ciphers suitable for widespread adoption. HC-256 is a stream cipher, designed to provide bulk encryption in software at high speeds while permitting strong confidence in its security. A 128-bit variant was submitted in 2004 as an eSTREAM cipher candidate; it has been selected as one of the four final contestants in the software profile [2,4] in 2008 as the most advanced scheme for stream cipher applications in advanced network environment.</p></sec><sec id="s1_3"><title>1.3. Large Noncoding DNA &amp; RNA</title><p>In relation to DNA analysis, visualization methods play a key role in the Human Genome Project (HGP) [<xref ref-type="bibr" rid="scirp.38482-ref19">19</xref>]. After HGP completed successfully, a public research consortium—the Encyclopedia of DNA Elements (ENCODE) was launched by the National Human Genome Research Institute (NHGRI) in 2003 to find all functional elements in the human genome as one of the most critical projects by NHGRI to explore genomes after HGP.</p><p>In 2012, ENCODE released a coordinated set of 30 papers published in key Journals of Nature, Genome Biology and Genome Research. These publications show that approximately 20% of noncoding DNA in the human genome is functional while an additional 60% is transcribed with no known function [<xref ref-type="bibr" rid="scirp.38482-ref20">20</xref>]. Much of this functional non-coding DNA is involved in the regulation of the expression of coding genes [<xref ref-type="bibr" rid="scirp.38482-ref21">21</xref>]. Furthermore, the expression of each coding gene is controlled by multiple regulatory sites located both near and distant from the gene. These results demonstrate that gene regulation is far more complex than was previously believed [<xref ref-type="bibr" rid="scirp.38482-ref22">22</xref>]. Mammalian genomes encode thousands of large noncoding RNAs (lncRNAs), many of which regulate gene expression, interact with chromatin regulatory complexes, and are thought to play a role in localizing these complexes to target loci across the genome [<xref ref-type="bibr" rid="scirp.38482-ref23">23</xref>]. Associated with different international projects, larger numbers of Genome Databases are established and mass Genomewide gene expression measurements are developed.</p><p>Due to huge amount of DNA sample collections and extremely difficulties to determine their variation properties in wider applications [24-30], it is essential for us to extend advanced DNA analysis models, methods and tools in further extensions to explore emerging models and concepts to interpret complex interactions among complicated sets of DNA sequences in real environments.</p></sec><sec id="s1_4"><title>1.4. DNA Analysis</title><p>DNA analysis plays a key role in modern genomic application [<xref ref-type="bibr" rid="scirp.38482-ref19">19</xref>]. The HGP is heavily relevant to advanced DNA sequencing and analysis techniques. DNA sequences are composed of four Meta symbols on {A, T, G, C} as basic structure. Classical DNA double helix structure makes the first level of pair construction of DNA sequences with A &amp; T and G &amp; C complementary structures as the first level of symmetric relationships. A typical DNA sequencing result is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>(a). Four Meta symbols could be separated as four projective sequences.</p><p>In ENCODE, recent Genomic analysis results are indicated that encoded sequences have only 20 percent in human genomes and around 80 percent genomes look like useless sequences. Under further assumptions, it seems that additional symmetric properties are required to satisfy the second, third and higher levels of structural constructions to explore complex interactive properties [24-30].</p><p>In current situation, it is necessary for advanced researchers to shift targets in computational cell biology from directly collecting sequential data to making higherlevel interpretation and exploring efficient content-based retrieval mechanism for genomes. Using higher dimensional visualization tools, their complex interactive properties could be organized as different visual maps systematically.</p></sec><sec id="s1_5"><title>1.5. Variant Construction and DNA</title><p>Variant construction is a new structure composed of logic, measurement and visualization models to analyze</p><p>0 - 1 sequences under variant conditions. The further details of this construction can be checked on variant logic [31,32], 2D maps [33,34], variant pseudo-random number generator [35-37], DNA maps [<xref ref-type="bibr" rid="scirp.38482-ref38">38</xref>] and variant phase spaces [<xref ref-type="bibr" rid="scirp.38482-ref34">34</xref>]. Since the variant system uses another set of four Meta symbols <img src="2-5300525\d8e7b309-4ba5-4d58-8949-1e0fe1f61f5f.jpg" /> to describe system, a typical correspondence shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>(b) may provide a natural mapping between DNA and variant data sequences.</p><p>Since DNA sequences are played an essential role to explore different symmetric properties based on analysis approaches, in this paper, measurement and visual models are proposed systematically to use a fixed segment structure to measure four Meta symbols distributions in their spectrum construction. Under this construction, refined symmetric features can be identified from various polarized distributions and further symmetric properties are visualized.</p></sec><sec id="s1_6"><title>1.6. Target of This Paper</title><p>The target of this paper is to establish the Variant Map System (VMS) as a unified framework to analyze complex DNA interactions on both artificial and natural DNA sequences. The VMS has designed to use variant logic schemes [31-38] applying multiple maps on four Meta symbols as DNA or RNA representations. System architecture of key components and core mechanism on the VMS are described. Key modules, equations and their I/O parameters are discussed. Applying the VM System, two sets of real DNA sequences from both human (noncoding DNA) and corn (coding DNA) genomes are collected in comparison with pseudo DNA sequences generated artificially by HC-256 to show their intrinsic properties in higher levels of similar relationships among DNA sequences on 2D maps. Further descriptions and discussions are provided respectively.</p></sec></sec><sec id="s2"><title>2. System Architecture</title><p>In this section, system architecture and their core components are discussed with the use of diagrams. The refined definitions and equations of this system are described in the next section—Variant Map System.</p><sec id="s2_1"><title>2.1. Architecture</title><p>T Architecture he four components of a variant map system are the Binary To DNA (BTD), the Binary Probability Measurement (BPM), the Mapping Position (MP), and the Visual Map (VM) as shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>The architecture is shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(a) with the key modules of the four core components being shown in Figures 2(b)-(e) respectively.</p><p>In the first part of the system, the t-th sequence <img src="2-5300525\7886c8e7-08d4-4c5b-84da-ccfa62840cae.jpg" /> on either {0, 1} or {A, G, T, C} are input data to get into the BTD module. The main function of the BTM is to output a unified sequence <img src="2-5300525\031a9c62-60ef-4d5a-95d9-92901db328f6.jpg" /> either to transfer a 0 - 1 sequence or to keep a DNA sequence as a pseudo or pure DNA sequence under a set of controlled parameters.</p><p>Using this unified DNA sequence, four vectors of probability measurements are created from the t-th selected DNA sequence with <img src="2-5300525\11002fc1-c5b3-4884-866c-3d8f4ec21905.jpg" /> elements as an input. Multiple segments are partitioned by a fixed number of n elements for each segment; at least <img src="2-5300525\b3c1dcb1-2dd9-4fe0-b192-8d5351fefefa.jpg" /> segments can be identified by the BPM component. Next component uses the four vectors of probability measurements and a given k value as input data, a pair of position values are created for each Meta symbol. Four pairs of values are generated by the MP component. Then, in order to process multiple selected DNA sequences, all selected sequences are processed by the VM component and each sequence may provide a set of pair values to generate relevant variant maps to indicate their distribution properties respectively.</p><p>With eight parameters in an input group, there are three sets of parameters in the intermediate group and one set of parameters in the output group.</p><p>The three groups of parameters are listed as follows.</p><sec id="s2_1_1"><title>Input Group:</title><p>t An integer indicates the t-th DNA sequence selected, <img src="2-5300525\7f576128-9691-43cc-a4ad-0e06d249c7bd.jpg" /></p><p><img src="2-5300525\80ca99d8-032e-40af-b961-7fd1e3ce089f.jpg" />An integer indicates a relationship distance among elements in a binary sequence, <img src="2-5300525\c8181a60-2f70-4e67-96c7-08a757d984db.jpg" /><img src="2-5300525\657c97a6-93ee-4846-afd6-1697b369f10a.jpg" />An integer indicates the mode of elements in a sequence, <img src="2-5300525\13d08030-eacf-48e4-89e6-64c8bd5ba379.jpg" />, <img src="2-5300525\a4f4d830-2e26-40f3-80e5-2326ccc35d95.jpg" />for a DNA sequence, <img src="2-5300525\1c4ff774-05f7-41c5-b055-16346aa8f32f.jpg" />for a binary sequence</p><p><img src="2-5300525\b875aa5d-b9e2-4b45-8806-b19f783fabaa.jpg" />An integer indicates the number of elements in the t-th DNA sequence, <img src="2-5300525\0cf68ab2-73b2-4f1e-b04d-153ed4451df8.jpg" /></p><p><img src="2-5300525\b5f4bbca-bbcb-44a6-a2b1-e33672db3a5b.jpg" />An input data vector with <img src="2-5300525\7f5303d3-aaf2-4096-aece-4bfeaaaa2426.jpg" /> elements, <img src="2-5300525\cd3f229c-83a4-4caa-910c-462c467fb101.jpg" /></p><p><img src="2-5300525\d34ca5dc-a4a8-4212-b0a1-36ea594774b3.jpg" />An integer indicates the number of elements in a segment, <img src="2-5300525\f6cd8ee7-f7b5-4c25-a933-288a42167e6c.jpg" /></p><p>V A symbol is selected from four DNA symbols <img src="2-5300525\fcf80f14-f55d-4605-9610-0ab22c429057.jpg" /></p><p><img src="2-5300525\9dfc4509-7473-4e62-8be9-f6703ab600d9.jpg" />An integer indicates the control parameter for mapping, <img src="2-5300525\33b6dbf5-d7ae-4cd5-87da-ca1f8f41f607.jpg" /></p></sec><sec id="s2_1_2"><title>Intermediate Group:</title><p><img src="2-5300525\ca87f1da-12f0-40bb-be21-0336d56f20e1.jpg" />A unified DNA vector with <img src="2-5300525\d5b3f677-53f5-4c07-800e-a3544e36cf01.jpg" /> elements, <img src="2-5300525\ad5e3d3a-1709-448c-8a01-93f019709b9a.jpg" /></p><p><img src="2-5300525\069a10ff-0c1a-4761-8157-96b5703760a2.jpg" />Four sets of probability measurements with <img src="2-5300525\ace4e482-9d1a-4dab-8879-fec37a8befd1.jpg" /></p><p><img src="2-5300525\2fc6e395-ada5-4594-9983-5ab4fdaacc2c.jpg" />Four paired values, <img src="2-5300525\810ddc90-56ff-46fb-a612-872352b2d0e1.jpg" /></p></sec><sec id="s2_1_3"><title>Output Group:</title><p><img src="2-5300525\d2a3b912-3440-4281-b04e-84cdd9706998.jpg" />Four 2D maps, <img src="2-5300525\de2015ab-e9e4-4891-ab6c-6f82c1e1d758.jpg" /></p></sec></sec><sec id="s2_2"><title>2.2. BTD Binary to DNA</title><p>The BTD component shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(b) is composed of one module: BTD itself. Five parameters are shown as</p><p>input signals and one unified vector is generated by the BTD component as the output group.</p><sec id="s2_2_1"><title>Input Group:</title><p>t An integer indicates the t-th DNA sequence selected, <img src="2-5300525\ae8c1cf5-f2c0-44ec-ae73-6b5e0c2479b6.jpg" /></p><p><img src="2-5300525\4cbf41d6-129b-480b-bd9b-a2c452859d08.jpg" />An integer indicates a relationship distance among elements in a binary sequence, <img src="2-5300525\397c2a6a-5bfb-410c-a80b-f3469ae883a9.jpg" />mode An integer indicates the mode of elements in a sequence, <img src="2-5300525\a0ed5852-7deb-4b3f-907a-67b93f5a9f07.jpg" />, <img src="2-5300525\fd4732a3-e129-4e46-bcdc-03fcd336618c.jpg" />for a DNA sequence, mode = 1 for a binary sequence</p><p><img src="2-5300525\280aeb5f-0e9a-4d16-ad6b-8c35e040dbdc.jpg" />An integer indicates the number of elements in the t-th DNA sequence, <img src="2-5300525\3dfd38f7-ebf3-4e4d-9c2e-ebc9427c0532.jpg" /></p><p><img src="2-5300525\515088b4-87f0-4d94-a9e7-e7342cacd15b.jpg" />An input data vector with <img src="2-5300525\83e01daa-4a3c-4ada-9e4b-7765d06f4c43.jpg" /> elements, <img src="2-5300525\99293e6f-d8c7-48b6-82b7-34f251773a78.jpg" /></p></sec><sec id="s2_2_2"><title>Output Group:</title><p><img src="2-5300525\3a47d011-4580-4e50-a2a9-b7306622ce4a.jpg" />A unified data vector with <img src="2-5300525\943940ad-966b-4ec0-89b7-e49e214905a6.jpg" /> elements, <img src="2-5300525\f734f330-08b8-4b4c-83e4-1beeb6415adf.jpg" /></p><p>The BTD component uses an input vector on either binary or DNA format as input, under a set of input parameters to process transformation. The output of the BTD component is composed of a unified vector of DNA format in a given condition.</p></sec></sec><sec id="s2_3"><title>2.3. BPM Binary Probability Measurement</title><p>The BPM component shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(c) is composed of two modules: BM Binary Measure and PM Probability Measurement. Three parameters are listed as input signals; four vectors of binary measures are outputted from the BM component as an intermediate group and four sets of probability measurements are outputted as an output group.</p><sec id="s2_3_1"><title>Input Group:</title><p><img src="2-5300525\78c6d4da-20a9-4d55-80e8-91722c68fa16.jpg" />An integer indicates the number of elements in a segment, <img src="2-5300525\a714cd3c-2a19-4f16-a343-ebe488320e6f.jpg" /></p><p>V A symbol is selected from four DNA symbols, <img src="2-5300525\724611f9-7828-4299-b928-78063a1ae72f.jpg" /></p><p><img src="2-5300525\62d8d1f9-1a02-4c29-869b-a59a982b29d7.jpg" />A DNA vector with <img src="2-5300525\0226dc38-8890-4efe-9841-b57da08a7804.jpg" /> elements, <img src="2-5300525\96e5d7a1-2078-4b83-a4b5-1dc2d2a32096.jpg" /></p></sec><sec id="s2_3_2"><title>Intermediate Group:</title><p><img src="2-5300525\41edbe96-f2b3-48d0-a123-2b67ac846ef7.jpg" />Four 0 - 1 vectors with <img src="2-5300525\a57a0394-e180-4601-a1c9-30a75478c861.jpg" /> elements, <img src="2-5300525\d3b85c1e-d81b-42a1-adc1-355209733228.jpg" /></p></sec><sec id="s2_3_3"><title>Output Group:</title><p><img src="2-5300525\73d82321-9297-4f5c-a8d1-3934e9cc7d9f.jpg" />Four sets of probability measurements with <img src="2-5300525\f571843a-01b9-4846-8818-10be5595165e.jpg" /></p><p>The BPM component transforms a selected DNA sequence to generate four 0 - 1 vectors by BM module for the input DNA sequence. Then four probability vectors are generated by the PM module as the output of the BPM under a fixed length of segment condition.</p></sec></sec><sec id="s2_4"><title>2.4. MP Mapping Position</title><p>The MP component shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(d) is composed of three modules: HIS Histogram, NH Normalized Histogram and PP Pair Position. Two parameters are listed as input signals; four histograms and four normalized histograms are generated from the HIS component and the NH component as intermediate groups respectively. Four paired values are generated by the PP component as the output group.</p><sec id="s2_4_1"><title>Input Group:</title><p><img src="2-5300525\70c7a6a1-65a9-4f5a-a49d-712aeb7771c1.jpg" />Four sets of probability measurements with <img src="2-5300525\b816229a-9609-4843-ada4-86882b63de68.jpg" /></p><p><img src="2-5300525\e2a6adac-1776-48d7-b78d-4b09223cee77.jpg" />An integer indicates the control parameter for mapping, <img src="2-5300525\31d4997c-8a2f-4233-8fbc-b0e6935ac90a.jpg" /></p></sec><sec id="s2_4_2"><title>Intermediate Group:</title><p><img src="2-5300525\f8e86435-832e-4b10-8ef0-0aa72a9a0d33.jpg" />Four histograms for relevant probability measurements, <img src="2-5300525\24f6aa00-1c62-4393-b09a-8cbb30c8a709.jpg" /></p><p><img src="2-5300525\c556ffe4-92a2-4e50-a5b2-49084810567c.jpg" />Four normalized histograms for relevant probability measurements, <img src="2-5300525\cc281a21-2486-4c43-a353-c0e7a0d80314.jpg" /></p></sec><sec id="s2_4_3"><title>Output Group:</title><p><img src="2-5300525\4d68d8d0-fb44-436f-8650-33fe8f944174.jpg" />Four paired values, <img src="2-5300525\6d523184-9b7a-48e3-98d1-624dcae387db.jpg" /></p><p>The MP component uses probability measurements as input, under a given k condition to generate each relevant histogram and its normalized distribution. The output of the MP component is composed of four paired values controlled in a given condition.</p></sec></sec><sec id="s2_5"><title>2.5. VM Visual Map</title><p>The VM component shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(e) is composed of one module: VM Visual Map. Three parameters are input signals. Collected all selected DNA sequences, four 2D maps are generated by the VM component as the output result.</p><sec id="s2_5_1"><title>Input Group:</title><p><img src="2-5300525\baa83cb5-a6a1-465a-931f-7857306fbad4.jpg" />All DNA sequences are selected, <img src="2-5300525\50f152ad-40b2-4ceb-b348-179ed2977e0e.jpg" /></p><p><img src="2-5300525\8ab5a9e9-6fe5-42ca-8b03-c3c19d694f19.jpg" />An input data vector with <img src="2-5300525\cff19908-e333-4f57-b3cd-7b87946460ad.jpg" /> elements, <img src="2-5300525\4e435828-a8bd-4cce-be2c-2aa5b8ff8119.jpg" /></p><p><img src="2-5300525\e359e45f-f430-40d2-bbe9-824e59482048.jpg" />Four paired values for the t-th DNA sequence, <img src="2-5300525\003751b8-d483-4260-9494-0c3d7e756cf8.jpg" /></p></sec><sec id="s2_5_2"><title>Output Group:</title><p><img src="2-5300525\6ef0a2f4-a2a0-431a-a529-2ff408500dd6.jpg" />Four 2D maps, <img src="2-5300525\8362384f-b11e-4fff-a850-7c49d90487fb.jpg" /></p><p>The VM component processes all selected DNA sequences as input to generate paired values for each sequence. The output of the VM component is composed of four 2D maps to show the final visual distribution for the system.</p></sec></sec></sec><sec id="s3"><title>3. Variant Map System</title><sec id="s3_1"><title>3.1. Initial Preparation</title><p>Let r an input parameter make all pairs of elements with r distance in a binary sequence to be a pseudo DNA vector, mode a controlled parameter indicate various pairs of operations performed if<img src="2-5300525\e70081f9-81de-4880-8e42-069ba2d3decb.jpg" />. Denote <img src="2-5300525\02772c5c-5d1c-4d11-90d9-0422b9230dd5.jpg" /> a binary base and <img src="2-5300525\b8082342-93c3-48db-a957-ac2b37e6792b.jpg" /> a DNA base respectively.</p></sec><sec id="s3_2"><title>3.2. BTD Module</title><p>Let <img src="2-5300525\03c5833d-d3b3-4306-b5c6-62fbe9ee1192.jpg" /> an input sequence with N elements,<img src="2-5300525\9b315bf3-51cd-414f-bcb0-517664f4d630.jpg" /> ,</p><p><img src="2-5300525\84b168ee-987f-4bb5-b071-b4d396fbf56c.jpg" />. This input vector could be expressed as follows.</p><p><img src="2-5300525\0b493d0d-8e5f-475e-89fc-0ad442b152f2.jpg" />,</p><disp-formula id="scirp.38482-formula59814"><label>. (1)</label><graphic position="anchor" xlink:href="2-5300525\eb2239df-72f7-40b4-833a-449400426bc3.jpg"  xlink:type="simple"/></disp-formula><p>Let X denote a DNA sequence with N elements, D denote a symbol set with four elements i.e.<img src="2-5300525\57a7d51f-ce83-48d3-a0ad-eb96f8c6072b.jpg" />. This type of a DNA sequence can be described by a four valued vector as follows:</p><p><img src="2-5300525\6b920f8f-3d09-4111-93ce-0fb4f7d52738.jpg" />,</p><disp-formula id="scirp.38482-formula59815"><label>. (2)</label><graphic position="anchor" xlink:href="2-5300525\8ccd87f1-bcee-4356-b200-17d71687d092.jpg"  xlink:type="simple"/></disp-formula><p>From this input and associated parameters, following operations are performed.</p><p>If<img src="2-5300525\6edd3906-2fd3-4dc2-b7ba-050b899e38b4.jpg" />, for all<img src="2-5300525\b7c89dc6-ddfc-48f5-ae2b-69bbe5e626c3.jpg" />, the output vector is equal to the input vector.</p><disp-formula id="scirp.38482-formula59816"><label>(3)</label><graphic position="anchor" xlink:href="2-5300525\1867ebe6-fa8c-47dd-b9d6-fba8ba43a3cf.jpg"  xlink:type="simple"/></disp-formula><p>If<img src="2-5300525\df34413a-f0ac-4601-8928-2b7ef9adf7f8.jpg" />, for all pairs of I and <img src="2-5300525\eb08ddfa-c94e-4d3f-8e29-20892837c00b.jpg" /> elements of <img src="2-5300525\acb214f6-6b90-4839-b617-a58fcc1c3c11.jpg" /> the I-th output element <img src="2-5300525\3ef27511-6d16-4060-8357-9d0ea90d722a.jpg" /> can be determined by the corresponding conditions shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>(b) as follows.</p><p><img src="2-5300525\5c63e7ec-b69f-4a9c-b535-73714140b580.jpg" /></p><disp-formula id="scirp.38482-formula59817"><label>(4)</label><graphic position="anchor" xlink:href="2-5300525\8e6531cf-5419-4add-83a0-e75a1d02b488.jpg"  xlink:type="simple"/></disp-formula><p>In both conditions, <img src="2-5300525\12f2b1d0-2fc7-4bb8-b2b3-c71b6cb4ff4a.jpg" />will be a unified vector with four values as the output of the BTD shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(b).</p><p>e.g. Let a binary sequence<img src="2-5300525\c02cd2a3-fbbd-42b5-a366-cb4c53d51f6f.jpg" />, three pseudo DNA sequences <img src="2-5300525\a0d540d9-f7ee-45ed-947e-4f673f63b57b.jpg" /> can be represented as follows.</p><p><img src="2-5300525\f8f35f65-d5b8-448f-8cf5-03fe46df17ce.jpg" /></p><p>Selecting a certain <img src="2-5300525\a92a6616-dc75-4417-9167-4b1d73942e71.jpg" /> value, a relevant pseudo DNA sequence can be generated from an input binary sequence.</p></sec><sec id="s3_3"><title>3.3. BM Module</title><p>For a given I-th element, four projective operators can be defined and denoted as <img src="2-5300525\593abee9-c413-4fef-8e43-2857b12fc838.jpg" />.</p><disp-formula id="scirp.38482-formula59818"><label>(5)</label><graphic position="anchor" xlink:href="2-5300525\2439d851-e192-41c2-9860-9fe6498df97c.jpg"  xlink:type="simple"/></disp-formula><p>Applying the four operators to all elements, the DNA sequence X can be reorganized into the four binary sequences of 0 - 1 values. i.e.</p><p><img src="2-5300525\a91c1313-cd61-465c-b477-1288523f420c.jpg" /></p><disp-formula id="scirp.38482-formula59819"><label>(6)</label><graphic position="anchor" xlink:href="2-5300525\af4503ca-2a68-4d45-9a6c-7943d6eadeca.jpg"  xlink:type="simple"/></disp-formula><p>e.g. let a DNA sequence <img src="2-5300525\a0ad5fd3-da90-433f-bcea-12c9310ac840.jpg" />, its four binary sequences can be represented as follows.</p><p><img src="2-5300525\bb5f38e5-b38d-4215-a362-f8aa96dada46.jpg" /></p><p>It is interesting to notice that the basic relationship between a DNA sequence X and its four <img src="2-5300525\24edf5e7-1437-4cd0-8c23-2469afd6347d.jpg" /> sequences are exactly same as in a modern DNA sequencing procedure to separate a selected DNA sequence into the four Meta symbol sequences shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>(a). This correspondence could be the key feature to apply the proposed scheme naturally in simulating complex behaviors for any DNA sequence.</p><p>The projection <img src="2-5300525\535a2044-17d4-43e6-b1af-3f8fe7218ebe.jpg" /> provides the essential operation in the BM component as the first module shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(c).</p></sec><sec id="s3_4"><title>3.4. PM Module</title><p>For this set of the four binary sequences, it is convenient to partition them into m segments and each segment contained a fixed number of n elements.</p><p>For the l-th segment, let <img src="2-5300525\c66c80eb-990d-48d0-ad8b-4a28890a9654.jpg" /> , the I-th position will be<img src="2-5300525\5627e696-73d6-4904-98d7-7fda373bad17.jpg" />, four probability measurements <img src="2-5300525\37aca0aa-a61d-4e39-9f6d-8a5aed18d80f.jpg" /> can be defined.</p><disp-formula id="scirp.38482-formula59820"><label>(7)</label><graphic position="anchor" xlink:href="2-5300525\b1dbfb1f-8c8a-4347-a5fb-742a2a333b8b.jpg"  xlink:type="simple"/></disp-formula><p>Under this construction, four sets of probability measurements established.</p><disp-formula id="scirp.38482-formula59821"><label>(8)</label><graphic position="anchor" xlink:href="2-5300525\99a3fbea-6467-4860-9e08-31f8665fa82a.jpg"  xlink:type="simple"/></disp-formula><p>The probability operator <img src="2-5300525\87cf5b31-c5d3-4538-b152-1a5af05e42a1.jpg" /> generates four probability measurement vectors in the PM component as the second module shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(c). After the BM and PM processes, the whole procedure of the BPM component is complete in <xref ref-type="fig" rid="fig2">Figure 2</xref>(c).</p></sec><sec id="s3_5"><title>3.5. HIS Module</title><p>Since the BPM generates four sets of probability measurement, it is necessary to perform further operations in the MP component shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(d) as follows.</p><p>In the HIS component as the first module in <xref ref-type="fig" rid="fig2">Figure 2</xref>(d), each probability sequence <img src="2-5300525\f1ee1c9f-58ca-4ea1-9d3d-e11581395595.jpg" /> can be calculated from n positions, at most n + 1 distinguished values identified in a vector. Under this organization, a histogram distribution can be established.</p><p>Let <img src="2-5300525\04a7b212-a315-4081-8a9d-ddbb6ebec771.jpg" /> be a histogram operator, for each position, it satisfies following relation,</p><disp-formula id="scirp.38482-formula59822"><label>(9)</label><graphic position="anchor" xlink:href="2-5300525\89e76cbd-6ace-420e-8023-1661a0e4ad32.jpg"  xlink:type="simple"/></disp-formula><p>Collecting all possible values, a histogram distribution can be established,</p><disp-formula id="scirp.38482-formula59823"><label>(10)</label><graphic position="anchor" xlink:href="2-5300525\e0a98286-52b8-4c94-bc01-c352f1103114.jpg"  xlink:type="simple"/></disp-formula><p>The histogram <img src="2-5300525\73590dc6-62f4-412f-81b0-2c276e7e30e5.jpg" /> is the output of the HIS module. Four histograms are generated after HIS process. Further normalized process will be performed in the NH component as the second module in <xref ref-type="fig" rid="fig2">Figure 2</xref>(d).</p></sec><sec id="s3_6"><title>3.6. NH Module</title><p>Under this construction, a normalized histogram can be defined as</p><disp-formula id="scirp.38482-formula59824"><label>(11)</label><graphic position="anchor" xlink:href="2-5300525\58964a49-9f1c-42d4-9e0a-395854e56d1e.jpg"  xlink:type="simple"/></disp-formula><p>After the NH component processed, its output provides the PP component for further operations as the third module in <xref ref-type="fig" rid="fig2">Figure 2</xref>(d).</p></sec><sec id="s3_7"><title>3.7. PP Module</title><p>Relevant probability vectors have <img src="2-5300525\91088096-233e-4ccd-ba1b-54d8992e11ca.jpg" /> distinguished values; four sets of normalized vectors can be organized as a linear order as follows,</p><disp-formula id="scirp.38482-formula59825"><label>(12)</label><graphic position="anchor" xlink:href="2-5300525\9c8ffe49-25bb-4ddd-b919-cd4687436ad8.jpg"  xlink:type="simple"/></disp-formula><p>Under this condition, four linear sets of probability vectors are established,</p><disp-formula id="scirp.38482-formula59826"><label>(13)</label><graphic position="anchor" xlink:href="2-5300525\1cff89a3-6a47-4c24-a883-c1a408c082da.jpg"  xlink:type="simple"/></disp-formula><p>For four vectors, their components can be normalized respectively,</p><disp-formula id="scirp.38482-formula59827"><label>(14)</label><graphic position="anchor" xlink:href="2-5300525\d9be6d18-80b0-4af7-ad04-f348f4fa2e35.jpg"  xlink:type="simple"/></disp-formula><p>Four sets of probability vectors are composed of a complete partition on their measurements.</p><p>Using this set of measurements, two mapping functions can be established to calculate a pair of values to map analyzed DNA sequence into a 2D map as follows.</p><p>Let <img src="2-5300525\36db421e-87b2-45ae-a421-577df1d4044d.jpg" /> and <img src="2-5300525\c3708470-c8a6-4f8c-a6ad-e58eeccd78e6.jpg" /> or <img src="2-5300525\1f12acdd-499a-4cd6-a7fb-82b5365e8913.jpg" /> be a pair of values defined by following equations,</p><p><img src="2-5300525\4540e6c8-7d58-4789-9c9a-90a78e59eac2.jpg" /></p><disp-formula id="scirp.38482-formula59828"><label>(15)</label><graphic position="anchor" xlink:href="2-5300525\3a51f192-a819-4985-88a9-d67185d96437.jpg"  xlink:type="simple"/></disp-formula><p>In the PP component, four paired values are generated and each pair indicates a specific position on a 2D map for the selected DNA sequence. The core operations of three key components: BTD, BPM and MP for a selected sequence are performed in Figures 2(b)-(d).</p></sec><sec id="s3_8"><title>3.8. VM Module</title><p>Since only one point of a 2D map is determined for a selected DNA sequence, it is essential to apply relative larger number of DNA sequences as inputs to generate visible distributions. This type of operations will be performed in the VM component shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>(e).</p><p>In a general condition, the VM component processes a selected data set <img src="2-5300525\438f1fe9-7c28-4611-803a-50cbd009068f.jpg" /> composed of T sequences, the t-th sequence with <img src="2-5300525\abb60d46-c03c-4089-b8d9-62c257a15da9.jpg" /> elements can be expressed by</p><p><img src="2-5300525\267d3e84-d4f7-4920-b9a5-efcff8e5aac0.jpg" /></p><p>Each sequence can be processed to apply the same procedures of the BTD, BPM and MP components. Since for each segment, its length n will be fixed for all selected sequences, it is essential to make number of segments be <img src="2-5300525\959035ff-4cea-47cc-bdb7-edac2b627331.jpg" /> in convention to match each sequence. Under this expression, the last module VM collects all T pairs of positions on relevant 2D visual maps as follows,</p><disp-formula id="scirp.38482-formula59829"><label>(16)</label><graphic position="anchor" xlink:href="2-5300525\0f7b10d6-ac15-4343-ab44-552dd7096fec.jpg"  xlink:type="simple"/></disp-formula><p>A sample 2D map of VM is shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>; this provides an assistant illustration for this type of visual maps on a case of multiple sequences.</p><p>Under this construction, a total number of T DNA sequences are transformed as T visual points on four 2D visual maps that would be help analyzers to explore their intrinsic symmetry properties among four binary sequences.</p></sec></sec><sec id="s4"><title>4. Sample Results on 2D Maps</title><p>Two types of data sets are selected for comparison. The first type of data sets is real DNA data sequence collected from both human and plan genomes to illustrate their differences on 2D maps. The second type of data set is collected from the Stream Cipher HC-256 to generate a pseudo random binary sequence under a certain condition.</p><sec id="s4_1"><title>4.1. DNA Data Resources</title><p>It is important to use some real DNA sequences to illustrate various test results of the VMS. Two sets of DNA sequences are selected and relevant resource features are described as follows.</p><p>The first data set originally comes from the human genome assembly version 37 and was taken from the reference sequences of 13 anonymous volunteers from Buffalo, New York. Hi-C technology used to analyze chromatin interaction role in genome. From a genomic analysis viewpoint, this set of data may contain more complex secondary or higher level structures. A special structure nearly the GRCh37 DNA sequence has been identified to explore their spatial characteristics. After positive and negative sequencing, each data file contain 2700 DNA sequences and each sequence has around 500 elements stored in two files left and right respectively.</p><p>The second DNA data set are selected from some plant gene database for comparison. One set of DNA sequences of Corn genomes are stored in file 201-500 that contains 2700 DNA sequences and each sequence has around 200 - 600 elements. It may be ordinary single sequences without complex secondary structures.</p></sec><sec id="s4_2"><title>4.2. Pseudo DNA Data Resources</title><p>The Stream Cipher HC-256 has being used to generate a binary sequence on a total length of 2700 &#215; 500 bits in the file hc256 that has been partitioned as 2700 subsequences and each sub-sequence in 500 bits.</p><p>Using the VMS in various parameters, three sets of pseudo DNA sequences are generated and their 2D maps are illustrated, analyzed and compared in following subsections.</p></sec><sec id="s4_3"><title>4.3. Sample Results</title><p>Using the three files of DNA sequences and one pseudo binary sequence in three parameters, six sets of 2D maps are listed in Figures 4-9 under different conditions to illustrate their spatial distributions using the VMS in a controllable environment.</p><p>In <xref ref-type="fig" rid="fig4">Figure 4</xref>, three groups of eighteen 2D maps are shown in the range of <img src="2-5300525\57681f60-e569-427d-838a-80227313b07c.jpg" /> <img src="2-5300525\113b61a4-b9bf-4b35-bf59-71e1f7527b60.jpg" /></p><p><img src="2-5300525\860f87e1-35d6-44c6-9be9-1e42697f4340.jpg" />for comparison; (a1)-(a6) six</p><p><img src="2-5300525\d333e3c7-554f-4717-ab1f-bf7afc93179c.jpg" />maps for the file Right; (b1)-(b6) six <img src="2-5300525\14d495a8-c856-4c1b-b98a-4d17df5089cc.jpg" /> maps for the file 201-500; (c1)-(c6) six <img src="2-5300525\1eec92d2-5fbb-437a-ac93-8348fe56889e.jpg" /> maps for the file hc256 respectively.</p><p>In <xref ref-type="fig" rid="fig5">Figure 5</xref>, four groups of sixteen 2D maps for the file right are listed in the range of <img src="2-5300525\7bc71cb7-b08e-47bd-8d37-347412e1ffc0.jpg" /> <img src="2-5300525\bdf8ec03-70c4-4d66-be6c-a966132a57bf.jpg" />; (a) group (a1)-(a4) four <img src="2-5300525\fe970f0b-1f35-43a5-8955-5c251f32f8ba.jpg" /> maps; (b) group (b1)-(b4) four <img src="2-5300525\ed3ba1e4-863c-4900-a7b9-ece064752d7e.jpg" /> maps; (c) group (c1)-(c4) four <img src="2-5300525\c6d0d54a-7e9b-487e-8f8c-9a43cf311869.jpg" /> maps; (d) group (d1)-(d4) four <img src="2-5300525\a8b907c8-8a39-4837-95e1-77076d9447ca.jpg" /> maps.</p><p>In <xref ref-type="fig" rid="fig6">Figure 6</xref>, four groups of sixteen 2D maps for the file hc256 are listed in the range of<img src="2-5300525\f761b060-ee41-4985-a911-d8b3298c5203.jpg" />,<img src="2-5300525\06b7f181-3e3b-4120-8e6f-d5bc303b8f71.jpg" />; (a) group (a1)-(a4) four <img src="2-5300525\4b75a266-6162-4491-8c83-dcf32c028f6c.jpg" /> maps; (b) group (b1)-(b4) four <img src="2-5300525\f379dd9b-0c9e-44de-90b6-2eb74aaf4b9f.jpg" /> maps; (c) group (c1)-(c4) four <img src="2-5300525\76aca83f-62aa-4218-8615-39393ecc42d7.jpg" /> maps; (d) group (d1)-(d4) four <img src="2-5300525\caec5b1d-f6af-470e-a9db-6e7b4be74240.jpg" /> maps.</p><p>In <xref ref-type="fig" rid="fig7">Figure 7</xref>, four groups of sixteen 2D maps for the file right are selected in the range of <img src="2-5300525\7d42abd3-fc13-449a-a65a-25a53e838f73.jpg" /> <img src="2-5300525\21cc3388-62e9-44c5-8c7c-d91ef85eeed4.jpg" />; (a) group (a1)-(a4) four <img src="2-5300525\6764f9ae-0023-48a3-b299-d67b27a08f28.jpg" /> maps; (b) group (b1)-(b4) four <img src="2-5300525\5bd77781-e3d0-4b28-bf63-6ef9bbb83c3b.jpg" /> maps; (c) group (c1)-(c4) four <img src="2-5300525\19905fad-19e1-42e2-a725-2f711c894369.jpg" /> maps; (d) group (d1)-(d4) four <img src="2-5300525\53c795eb-0be5-43dc-95a4-e44ddbbb94b6.jpg" /> maps.</p><p>In <xref ref-type="fig" rid="fig8">Figure 8</xref>, three groups of twelve 2D maps for the file hc256 are compared in the range of<img src="2-5300525\aefad229-918e-42ae-bfcc-707b04d36c52.jpg" />. <img src="2-5300525\5861318d-77d6-40e3-93da-e0cf450132cd.jpg" />(a) group (a1)-(a4) four <img src="2-5300525\18b111b0-bcec-4d5e-8259-82f4ff063911.jpg" /> maps<img src="2-5300525\51bd83b9-f17f-4009-a6f1-372126d6634d.jpg" />; (b) group (b1)-(b4) four <img src="2-5300525\52a90966-f308-4f81-a803-710c745e64e9.jpg" /> maps<img src="2-5300525\3f046ec0-e136-4510-b243-68523a4978fa.jpg" />; (c) group (c1)-(c4) four<img src="2-5300525\d25c6a15-3a09-417d-9d7f-a9da91286e5d.jpg" /> maps<img src="2-5300525\1efded12-8d6c-482e-9eea-832d4f4c3c16.jpg" />.</p><p>In <xref ref-type="fig" rid="fig9">Figure 9</xref>, three groups of twelve 2D maps for two files right and hc256 are compared in the range of<img src="2-5300525\b6f93de9-ff8a-43e2-8307-cd822a1e1397.jpg" />; (a) the file right n=15, mode=0; (b) the file hc256 n = 12, mode = 1, r = 1; (c) the file hc256 n = 12, mode = 1, r = 3; (a1)-(c1) <img src="2-5300525\d10f9cb7-9ecc-44f2-be45-b7c0487dd4f4.jpg" />maps; (a2)-(c2) <img src="2-5300525\5868bbed-a8a4-4814-a7c1-7ac0dca13f86.jpg" />maps; (a3)-(c3) <img src="2-5300525\af0965e7-e9bf-4349-ace4-3fffd5cb4635.jpg" />maps; (a4)-(c4) <img src="2-5300525\a12cc853-27da-4b7a-bb67-3c5818268a48.jpg" />maps.</p></sec><sec id="s4_4"><title>4.4. Result Analysis of 2D Maps</title><p>Six groups of 2D maps contain different information, it is necessary to make a brief discussion on their important issues as follows.</p><p>The first group of results shown in <xref ref-type="fig" rid="fig4">Figure 4</xref> presents</p><p>three sets of eighteen 2D maps from three data files: right, 201-500 and hc256 undertaken various lengths of basic segment from 3-50 to illustrate their variations respectively. Six 2D maps of each group in <xref ref-type="fig" rid="fig4">Figure 4</xref> (a1)-(a6) show significant trace on their visual distributions; the numbers of main visible clusters identified are decreased when the length of segment has being increased e.g. (a3)-(a6). However lesser length of segment does not provide refined visual distinctions with larger region in fuzzy areas e.g. (a1) and (a2). From a structural viewpoint, middle ranged numbers of length provide better clustering results e.g. (a3)-(a5) for further analysis targets. To check another six 2D maps of <xref ref-type="fig" rid="fig4">Figure 4</xref> (b1)-(b6) for the file 201-500, significantly different visual distributions can be observed than (a1)-(a6); the numbers of main visible clusters identified are decreased when the length of segment has being increased less significantly e.g. (b4)-(b6). However lesser length of segment does not provide refined visual distinctions with wider regions in fuzzy areas e.g. (b1)-(b3). In generalmiddle ranged numbers of length still provide better clustering effects e.g. (b4)-(b6) for further analysis purpose. To check six 2D maps of <xref ref-type="fig" rid="fig4">Figure 4</xref> (c1)-(c6) for the file hc256 r = 1, similar visual distributions can be observed than (a1)-(a6) and significantly differences are observed than (b1)-(b6); the numbers of main visible clusters identified are decreased when the length of segment has being increased less significantly e.g. (c3)-(c6). However lesser length of segment does provide refined visual distinctions with regions in fuzzy areas e.g. (b1). In general, middle ranged numbers of length still provide better clustering effects e.g. (c2)-(c4) for further analysis purpose. From their distributions, groups (a) and (c) have shared much stronger similar properties than group (b).</p><p>It is interesting to observe different maps when control parameter k changed. Four groups of sixteen 2D maps for the file right are shown in <xref ref-type="fig" rid="fig5">Figure 5</xref> on the range of<img src="2-5300525\5fa11e7d-5c41-463c-9693-461e587bdfc2.jpg" />; four groups in (a)-(d) provide four maps to share the same other parameters with different k values. Checking visible clus-</p><p>ters in different maps, it is important to notice nearly same numbers of clusters identified in the same group, but different groups may contain significantly different numbers. Lesser k value (e.g. k = 2) makes a tighter distribution and larger k value (e.g. k = 7) takes better separation on the maps. Through k = 7 maps provide better separation effects, it is easy to observe their y axis values already in 10<sup>8</sup> range.</p><p>Four groups of sixteen 2D maps for the file hc256 are shown in <xref ref-type="fig" rid="fig6">Figure 6</xref> in the range of <img src="2-5300525\b488d075-5b2d-45f5-bc0b-cdbb167352fd.jpg" /> <img src="2-5300525\75597d12-b9d2-4a4e-9e24-21f40a436aea.jpg" />. This group of 2D maps can be compared with 2D maps in <xref ref-type="fig" rid="fig5">Figure 5</xref>. Under the same parameters, similar visible effects and feature clustering properties could be observed if various k values are selected.</p><p>Using a set of selected parameters, two groups of eight 2D maps are compared in <xref ref-type="fig" rid="fig7">Figure 7</xref> for two files: left, right to explore higher levels of symmetric properties for secondary or higher levels of structures potentially contained in DNA sequences. Selected parameters are in the range of<img src="2-5300525\62be8124-e0f9-4997-8364-f95a17679d4e.jpg" />. Group (a) provides four <img src="2-5300525\c56568ca-8779-4c0d-a813-b5d94c023078.jpg" /> maps (a1)-(a4) for the file left; group (b) uses four <img src="2-5300525\8f004988-f08f-4174-a5fb-4fa782e1c841.jpg" /> maps (b1)-(b4) for the file right.</p><p>In convenient description, let ~ be a similar operator, for groups (a) &amp; (b), four pairs of {(a1)~(b1), (a2)~(b2), (a3)~(b3), (a4)~(b4)} maps i.e. (left-A ~ right-A, left-T ~ right-T, left-G ~ right-G, left-C ~ right-C) have a stronger similar distribution between left &amp; right. In addition, only two clustering classes could be significantly identified as {(a1)~(a2)~(b1)~(b2), (a3)~(a4)~(b3)~(b4)} i.e. (left-A ~ right-A ~ left-T ~ right-T, left-G ~ right-G ~ left-C ~ right-C) respectively. This type of similar clustering distributions may strongly indicate eight maps with intrinsically higher levels of DNA sequences with extra A-T &amp; G-C pairs of symmetric relationships between two files: left &amp; right.</p><p>Using a set of selected parameters, three groups of twelve 2D maps are listed in <xref ref-type="fig" rid="fig8">Figure 8</xref> for the file hc256, r = {1, 2, 3} to explore properties for their higher levels of structures potentially contained in pseudo DNA sequences. Selected parameters are in the range of<img src="2-5300525\ff692519-c8e2-4f3a-a860-3506c48d7293.jpg" />. Group (a) provides four <img src="2-5300525\291e6e4f-22d5-4a3a-9a08-81a7222fb814.jpg" /> maps (a1)-(a4) for r = 1; group (b) uses four <img src="2-5300525\1faeb9fa-1f60-44ae-91c4-b5c82a7bebd4.jpg" /> maps (b1)-(b4) for r = 2 (c) uses four <img src="2-5300525\9c0b2b6f-6195-4c4f-b69f-6a6ef7dca5cf.jpg" /> maps (c1)-(c4) for r = 3. Using a similar operator, for groups (a)-(c), four pairs of {(a1)~(b1)~(c1), (a2)~(b2)~ (c2), (a3)~(b3)~(c3), (a4)~(b4)~(c4)} maps i.e. (A(r = 1)~A(r = 2)~A(r = 3), &#183;&#183;&#183;, C(r = 1)~C(r = 2)~C(r = 3)) have a stronger similar distribution among r = {1, 2, 3}. In addition, only two clustering classes could be significantly identified as {(a1)~(a2)~(b1)~(b2)~(c1)~(c2), (a3)~ (a4)~(b3)~(b4)~(c3)~(c4)} i.e. three maps are shown in (A~T, G~C) respectively.</p><p>In a convenient comparison, using a set of selected parameters, three groups of twelve 2D maps are compared in <xref ref-type="fig" rid="fig9">Figure 9</xref> for the files: right and hc256, r = {1, 3} to check their distribution properties contained in both DNA and created pseudo DNA sequences. Group (a) provides four <img src="2-5300525\2fe8a079-7ae2-4489-a292-59d14c50c7fe.jpg" /> maps (a1)-(a4) for the file right; groups (b) and (c) provide four <img src="2-5300525\f2b95d52-b868-49c9-96f5-adf73e38f8c1.jpg" /> maps (b1)-(b4) for hc256, r = 1 (c) and (c1)-(c4) for hc256, r = 3.</p><p>Using a weak similar operator ;, for groups (a)-(c), four pairs of {(a1);(b1)~(c1), (a2);(b2)~(c2), (a3)~ (b3)~(c3), (a4)~(b4)~(c4)} maps have a stronger similar distribution between r = {1,3} and a weak similar distribution on A &amp; T cases. In addition, only two clustering classes could be significantly identified as {(a1)~(a2); (b1)~(b2)~(c1)~(c2), (a3)~(a4)~(b3)~(b4)~(c3)~(c4)} i.e. three maps are strongly shown in relationships among (A~|;T, G~C) for different cases respectively.</p><p>In addition, this set of results illustrates directly visual comparisons with stronger similarity between DNA and pseudo DNA on VMS maps, their similarly clustering distributions may indicate those maps with comparable mechanism to express real DNA sequences with extra A-T &amp; G-C pairs of symmetric relationships in their higher levels of relationships applying the Stream Cipher mechanism.</p></sec></sec><sec id="s5"><title>5. Conclusions</title><p>This paper proposes architecture to support the Variant Map System. Using a binary random sequence as input, a set of special pseudo DNA sequences can be generated. Under variant measures, probability measurement and normalized histogram, a pair of values can be determined by a series of controlled parameters. Collecting relevant pairs on multiple DNA sequences, four 2D maps can be generated.</p><p>The main results of this paper provide the VMS architecture description in diagrams, main components, modules, expressions and important equations for the VMS. Core models and diagrams, sample results are illustrated to apply two types of data sets selected from real DNA sequences and generated from the pseudo random sequences from the Stream Cipher HC-256 for comparison under the VMS testing. After proper set of parameters selected, suitable visual distributions could be observed using the VMS. Results in Figures 4-9 provide useful evidences systematically to support proposed VMS useful in checking higher levels of symmetric/similar properties among complex DNA sequences in both natural and artificial environment.</p><p>This construction could provide useful insights to spatial information on complex DNA expressions especially on large encoding RNA/DNA construction via 2D maps to explore higher levels of complex interactive environments in near future.</p></sec><sec id="s6"><title>6. Acknowledgements</title><p>Thanks to the school of software Yunnan University, to the key laboratory of Yunnan software engineering and the key laboratory for Conservation and Utilization of Bio-resource for excellent working environment, to the Yunnan Advanced Overseas Scholar Project (W8110305), the Key R&amp;D project of Yunnan Higher Education Bureau (K1059178) and National Science Foundation of China (61362014) for financial supports to this project.</p></sec><sec id="s7"><title>REFERENCES</title></sec><sec id="s8"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.38482-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">ESTREAM Project. http://en.wikipedia.org/wiki/ESTREAM</mixed-citation></ref><ref id="scirp.38482-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">H. J. Wu, “Stream Cipher HC-256,” 2004.http://www.ecrypt.eu.org/stream/p3ciphers/hc/hc256_p3.pdf</mixed-citation></ref><ref id="scirp.38482-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">M. Santha and U. V. Vazirani, “Generating Quasi-Random Sequences from Slightly Random Sources,” Journal of Computer and System Sciences, Vol. 33, No. 1, 1986, pp. 75-87. http://dx.doi.org/10.1016/0022-0000(86)90044-9</mixed-citation></ref><ref id="scirp.38482-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">P. Goutam and M. Subhamoy, “RC4 Stream Cipher and Its Variants,” CRC Press, Boca Raton, 2012.</mixed-citation></ref><ref id="scirp.38482-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">M. Gude, “Concept for a High-Performance Random Number Generator Based on Physical Random Noise,” Frequenz, Vol. 39, No. 7-8, 1985, pp. 187-190.</mixed-citation></ref><ref id="scirp.38482-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">D. Eastlake, S. D. Crocker and J. I. Schiller, “Randomness Requirements for Security, RFC 1750,” 1994.</mixed-citation></ref><ref id="scirp.38482-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">C. Plumb, “Truly Random Numbers,” Dr. Dobbs Journal, Vol. 19, No. 13, 1994, pp. 113-115.</mixed-citation></ref><ref id="scirp.38482-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">G. B. Agnew, “Random Source for Cryptographic Systems,” Springer-Verlag, Berlin, 1988, pp. 77-81.</mixed-citation></ref><ref id="scirp.38482-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">A. Gehani, T. LaBean and J. Reif, “DNA-Based Cryptography,” DIMACS Series in Discrete Mathematica and Theoretical Computer Science, Vol. 54, 2000, pp. 233-249. http://www.cs.duke.edu/~reif/paper/DNAcrypt/DNA5.DNAcrypt.pdf</mixed-citation></ref><ref id="scirp.38482-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">B. E. Bernstein, E. Birney, I. Dunham, et al., “An Integrated Encyclopedia of DNA Elements in the Human Genome,” Nature, Vol. 489, No. 7414, 2012, pp. 57-74. http://dx.doi.org/10.1038/nature11247</mixed-citation></ref><ref id="scirp.38482-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">E. Pennisi, “Genomics. ENCODE Project Writes Eulogy for Junk DNA,” Science, Vol. 337, No. 6099, 2012, pp. 1159-1161. http://dx.doi.org/10.1126/science.337.6099.1159</mixed-citation></ref><ref id="scirp.38482-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">M. Schooniger and A. von Haeseler, “Simulating Efficiently the Evolution of DNA Sequences,” Computer Applications in the Biosciences, Vol. 11, No. 1, 1995, pp. 111-115.</mixed-citation></ref><ref id="scirp.38482-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">F. Piva and G. Principato, “RANDNA: A Random DNA Sequence Generator,” Silico Biology, Vol. 6, No. 3, 2006, pp. 253-258.</mixed-citation></ref><ref id="scirp.38482-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">C. M. Gearheart, B. Arazi and E. C. Rouchka, “DNA-Based Random Number Generation in Security Circuitry,” Biosystems, Vol. 100, No. 3, 2010, pp. 208-214. http://dx.doi.org/10.1016/j.biosystems.2010.03.005</mixed-citation></ref><ref id="scirp.38482-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">O. O. Babatunde, “On Pseudorandom Number Generation from Programmable and Computable Biomolecules: Deoxyribonucleic (DNA) as a Novel Pseudorandom Number Generator,” World Applied Programming, Vol. 1, No. 3, 2011, pp. 215-227.</mixed-citation></ref><ref id="scirp.38482-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">G. C. Sirakoulis, “Hybrid DNA Cellular Automata for Pseudorandom Number Generation,” 2012 International Conference on High Performance Computing and Simulation (HPCS), Madrid, 2-6 July 2012, pp. 238-244. http://dx.doi.org/10.1109/HPCSim.2012.6266918</mixed-citation></ref><ref id="scirp.38482-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Y. P. Zhang, Y. Zhu, Z. Wang and R. O. Sinnott, “Index-Based Symmetric DNA Encryption Algorithm, The 4th International Congress on Image and Signal Processing, Shanghai, 15-17 October 2011. http://dtl.unimelb.edu.au/researchfile287042.pdf</mixed-citation></ref><ref id="scirp.38482-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Y. P. Zhang, L. He and B. C. Fu, “Research on DNA Cryptography, Applied Cryptography and Network Security,” InTech Press, 2012. http://www.intechopen.com/books/applied-cryptography-and-network-security/research-on-dna-cryptography</mixed-citation></ref><ref id="scirp.38482-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">E. Lieberman-Aiden, et al., “Comprehensive Mapping of Long-Range Interactions Reveals Folding Principles of the Human Genome,” Science, Vol. 326, No. 5950, 2009. pp. 289-293. http://dx.doi.org/10.1126/science.1181369</mixed-citation></ref><ref id="scirp.38482-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">M. B. Gerstein, A. Kundaje, M. Hariharan, et al., “Architecture of the Human Regulatory Network Derived from ENCODE data,” Nature, Vol. 489, No. 7414, 2012, pp. 91-100. http://dx.doi.org/10.1038/nature11245</mixed-citation></ref><ref id="scirp.38482-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">W. F. Doolittle, “Is Junk DNA Bunk? A Critique of ENCODE,” Proceedings of the National Academy of Sciences, Vol. 110, No. 14, 2013, p. 5294.</mixed-citation></ref><ref id="scirp.38482-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">J. M. Engreitz, A. Pandya-Jones, P. McDonel, et al., “Large Noncoding RNAs Can Localize to Regulatory DNA Targets by Exploriting the 3D Architecture of the Genome,” 2013.</mixed-citation></ref><ref id="scirp.38482-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">K. Sakamoto, “Molecular Computation by DNA Hairpin Formation,” Science, Vol. 288, No. 5469, 2000, pp. 1223-1226. http://dx.doi.org/10.1126/science.288.5469.1223</mixed-citation></ref><ref id="scirp.38482-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">A. Arneodo, C. Vaillant. et al., “Multi-Scale Coding of Genomic Information: From DNA Sequence to Genome Structure and Function,” Physics Reports, Vol. 498, No. 2, 2011, pp. 45-188. http://dx.doi.org/10.1016/j.physrep.2010.10.001</mixed-citation></ref><ref id="scirp.38482-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">S. Engela, A. Alemany and N. Forns, “Folding and Unfolding of a Triple-Branch DNA Molecule with Four Conformational States,” Philosophical Magazine, Vol. 91, No. 13, 2011, pp. 2049-2065. http://dx.doi.org/10.1080/14786435.2011.557671</mixed-citation></ref><ref id="scirp.38482-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">J. M. Urquiza, I. Rojas, et al., “Method for Prediction of Protein-Protein Interactions in Yeast Using Genomics/ Proteomics Information and Feature Selection,” Neurocomputing, Vol. 74, No. 16, 2011, pp. 2683-2690. http://dx.doi.org/10.1016/j.neucom.2011.03.025</mixed-citation></ref><ref id="scirp.38482-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">H. Y. Zhang and X. Y. Liu. “A CLIQUE Algorithm Using DNA Computing Techniques Based on Closed-Circle DNA Sequences,” Biosystems, Vol. 105, No. 1, 2011, pp. 73-82. http://dx.doi.org/10.1016/j.biosystems.2011.03.004</mixed-citation></ref><ref id="scirp.38482-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">B. Banfai, H. Jia, J. Khatun, et al., “Long Noncoding RNAs Are Rarely Translated in Two Human Cell Lines,” Genome Research, Vol. 22, No. 9, 2012, pp. 1646-1657. http://dx.doi.org/10.1101/gr.134767.111</mixed-citation></ref><ref id="scirp.38482-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">J. S. Wang and M. Yan, “Numerical Methods in Bioinformatics,” Science Press, Beijing, 2013.</mixed-citation></ref><ref id="scirp.38482-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">N. A. Tchurikov, O. V. Kretova, D. M. Fedoseeva, et al., “DNA Double-Strand Breaks Coupled with PARP1 and HNRNPA2B1 Binding Sites Flank Coordinately Expressed Domains in Human Chromosomes,” PLoS Genetics, Vol. 9, No. 4, 2013, Article ID: e1003429. http://dx.doi.org/10.1371/journal.pgen.1003429</mixed-citation></ref><ref id="scirp.38482-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">J. Z. J. Zheng and C. H. Zheng, “A Framework to Express Variant and Invariant Functional Spaces for Binary Logic,” Frontier of Electrical and Electronic Engineering in China, Vol. 5, No. 2, 2010, pp. 163-172.http://dx.doi.org/10.1007/s11460-010-0011-4</mixed-citation></ref><ref id="scirp.38482-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">J. Zheng, C. Zheng and T. Kunii, “A Framework of Variant Logic Construction for Cellular Automata,” In: A. Salcido, Ed., Cellular Automata—Innovative Modelling for Science and Engineering, InTech Press, 2011, pp. 325-352. http://www.intechopen.com/chapters/20706</mixed-citation></ref><ref id="scirp.38482-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Q. P. Li and J. Zheng, “2D Spatial Distributions for Measures of Random Sequences Using Conjugate Maps,” Proceedings of the 11th Australian Information Warfare and Security Conference, Perth, 2010. http://ro.ecu.edu.au/isw/34</mixed-citation></ref><ref id="scirp.38482-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">J. Zheng, C. Zheng and T. Kunii, “Interactive Maps on Variant Phase Spaces—From Measurements Micro Ensembles to Ensemble Matrices on Statistical Mechanics of Particle Models,” In: A. Salcido, Ed., Emerging Application of Cellular Automata, InTech Press, 2013, pp. 113-196. http://dx.doi.org/10.5772/51635</mixed-citation></ref><ref id="scirp.38482-ref35"><label>35</label><mixed-citation publication-type="other" xlink:type="simple">J. Zheng, “Novel Pseudo-Random Number Generation Using Variant Logic Framework,” The 2nd International Cyber Resilience Conference, Perth, 1-2 August 2011, pp. 100-104.http://igneous.scis.ecu.edu.au/proceedings/2011/icr/zheng.pdf</mixed-citation></ref><ref id="scirp.38482-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">W. Z. Yang and J. Z. J. Zheng, “PRNG Based on Variant Logic,” 7th International ICST Conference on Communications and Networking in China (CHINACOM 2012), Kunming, 8-10 August 2012, pp. 202-205. http://www.computer.org/csdl/proceedings/chinacom/2012/2698/00/06417476-abs.html</mixed-citation></ref><ref id="scirp.38482-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">W. Z. Yang and J. Zheng, “Variant Pseudo-Random Number Generator,” Hakin9 Extra, No. 6, 2012, pp. 28-31. http://hakin9.org/hakin9-extra-62012/</mixed-citation></ref><ref id="scirp.38482-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">W.Q. Zhang and J. Zheng, “Randomness Measurement of Pseudorandom Sequence Using Different Generation Mechanisms and DNA Sequence,” Journal of Chengdu University of Information Technology, Vol. 27, No. 6, 2012, pp. 548-555.</mixed-citation></ref></ref-list></back></article>