hdu 1583 dna assembly

Problem DescriptionFarmer John has performed DNA sequencing on his prize milk-producing cow, Bessie DNA sequences are ordered lists (strings) containing the letters ✀A✀, ✀C✀, ✀G✀, and ✀T✀. As is usual for DNA sequencing, the results are a set of strings that are sequenced fragments of DNA, not entire DNA strings. A pair of strings like ✀GATTA✀ and ✀TACA✀ most probably represent the string ✀GATTACA✀ as the overlapping characters are merged, since they were probably duplicated in the sequencing process. Merging a pair of strings requires finding the greatest overlap between the two and then eliminating it as the two strings are concatenated together. Overlaps are between the end of one string and beginning of another string, NOT IN THE MIDDLE OF A STRING. By way of example, the strings ✀GATTACA✀ and ✀TTACA✀ overlap completely. On the other hand, the strings ✀GATTACA✀ and ✀TTA✀ have no overlap at all, since the matching characters of one appear in the middle of the other, not at one end or the other. Here are some examples of merging strings, including those with no overlap: GATTA + TACA -> GATTACA TACA + GATTA -> TACAGATTA TACA + ACA -> TACA TAC + TACA -> TACA ATAC + TACA -> ATACA TACA + ACAT -> TACATGiven a set of N (2 <= N <= 7) DNA sequences all of whose lengths are in the range 1..7, find and print length of the shortest possible sequence obtainable by repeatedly merging all N strings using the procedure described above. All strings must be merged into the resulting sequence.InputThe input consists of multiple test cases. Each test case :Line 1: A single integer N Lines 2..N+1: Each line contains a single DNA subsequenceEnd of file.OutputFor each pair of input output the length of the shortest possible string obtained by merging the subsequences. It is always possible – and required – to merge all the input strings to obtain this string.Sample Input4GATTATAGGATCGACGCATSample Output13HintHint Explanation of the sample: Such string is "CGCATCGATTAGG". 总是WAMY CODE:用C,不用C++
2026年09月28日 12:07
有1个网友回答
网友(1):

废话不多说,代码直接上
////////////////////////////
//Title:hdu 1583 dna assembly
//Code by:汇蓝鸟
//程序在windows7 VC6.0下编译通过并满足题目条件
//////////////////////////////
#include
#include
#include
int DNACom(char *DNA_A,char *DNA_B)//用于验证两DNA片段公用的碱基个数
{
int DNA_offset=0;
while(*(DNA_A+DNA_offset))
{
if(*(DNA_A+DNA_offset)==*(DNA_B+DNA_offset))
{
DNA_offset++;
}
else
{
DNA_A++;
DNA_offset=0;
}
}
return DNA_offset;
}
int GetDNAComMAX(char *DNA_A,char *DNA_B)
{
int n1=DNACom(DNA_A,DNA_B);
int n2=DNACom(DNA_B,DNA_A);
return n1>n2?n1:n2;
}
bool legDNA(char *DestDNA)//判断DNA合法性,虽然题目中没说,但还是写上好
{
while(*(DestDNA))
{
if(*DestDNA!='A'&&*DestDNA!='G'&&*DestDNA!='C'&&*DestDNA!='T')
return false;
DestDNA++;
}
return true;
}
void mergedDNA(char *DNA_A,char *DNA_B)//连接两个DNA片段
{
char DNA_Buf[1024];
memset(DNA_Buf,0,sizeof(DNA_Buf));
int Com1=DNACom(DNA_A,DNA_B);
int Com2=DNACom(DNA_B,DNA_A);
if(Com1>Com2)
{
strcpy(DNA_Buf,DNA_A);
strcat(DNA_Buf,DNA_B+Com1);
}
else
{
strcpy(DNA_Buf,DNA_B);
strcat(DNA_Buf,DNA_A+Com2);
}
strcpy(DNA_A,DNA_Buf);
}

void main()
{
int nDNA,MAXCom=0,DNANeedMerge1,DNANeedMerge2;
char DNASequences[255][1024];//最终决定不采用链表,用数组和谐
bool UsingFlag[255];
printf("请输入您要输入的DNA片段个数:");
scanf("%d",&nDNA);
printf("请输入DNA片段:\n");
for(int i=0;i {
scanf("%s",DNASequences[i]);
if(!legDNA(DNASequences[i]))
{
printf("\n输入的DNA不合法,重新输入\n");
i--;
continue;
}
}

for(i=1;i<=255;i++)
{
UsingFlag[i]=true;
}
//穷举实现公共碱基数目最大的先合成,合成结果放在数组靠前的那个数组中,最大支持1024个碱基
for(int z=0;z {
for(i=0;i for(int j=i+1;j<=nDNA-1;j++)
{
if(UsingFlag[i]&&UsingFlag[j])
if(MAXCom {
MAXCom=GetDNAComMAX(DNASequences[i],DNASequences[j]);
DNANeedMerge1=i;
DNANeedMerge2=j;
}
}
mergedDNA(DNASequences[DNANeedMerge1],DNASequences[DNANeedMerge2]);
UsingFlag[DNANeedMerge2]=false;
MAXCom=0;
}
printf("\n最后的DNA结果是%s\n最终合成长度%d\n",DNASequences[0],strlen(DNASequences[0]));//最终结果在DNASequences[0]当中
}
//乍看之下程序的局限性还是有的,就是内存占用会比较多,最大也就支持1024个碱基(要增多改改数字就行了),采用链表会好很多,我这人比较懒,数组随便凑活用吧