NCKU Research Team Shined at International Competition
Tainan, Taiwan, October 13th, 2010
Assistant Prof. Hung-Yu Kao of Department of Computer Science and Information Engineering at National Cheng Kung University (NCKU), Tainan, Taiwan, and his supervised doctoral student Chih-Hsuan Wei have shined at the 3rd Critical Assessment of Information Extraction in Biology (BioCreative) challenge in September, 2010, in Bethesda, United States, winning first place in the 5 assessments of the total 9 assessments of gene normalization task with their network model AINet.
The Critical Assessment of Information Extraction systems in Biology, a renowned competition in the field of biological text mining, is initiated by the Centro Nacional de Investigaciones Oncologicas (Spanish National Cancer Research Centre), which is led by Dr. Alfonso Valencia of Spanish National Bioinformatics Institute, Spain, in 2003.
The challenge covers a variety of important issues of contemporary biological text mining, and the results and methods of the winning teams will be invited to be published in renowned journals. For example, the results of the previous challenge were published in the 2010 Nature Biotechnology.
Assistant Prof. Hung-Yu Kao and his supervised doctoral student Chih-Hsuan Wei have participated in the task of gene normalization in the challenge and developed a capture mining system, which does not require training data, by using a method of inference network integrating probability and statistics, being the only team among the contestants which has adopted the non-machine learning method.
In text which has weak signal, they have enlarged the genetic identification evidence with an innovative statistic inference method and applied the method to extract the accurate gene numbers from the articles. This is an innovative and effective application of inference model on bioinformatics issues.
The evaluation rules of the challenge have adopted the TAP-K, Threshold Average Precision-K of National Center for Biotechnology Information, U.S., to assess 507 documents, and among them, 50 documents that are most difficult to extract are divided into two groups for assessment.
In the 9 assessments of the 3 groups, each setting 3 different K-values (K=5, K=10 and K=20), Associate Prof. Hung-Yu Kao and doctoral student Chih-Hsuan Wei have put on outstanding performance in the task of gene normalization, winning first place in 5 assessments of the total 9 assessments and also first place in the overall score.
Currently, the large biological knowledge base still relies on professional force extracting knowledge from the texts and the cost is extremely high.
The network model AINet developed by Associate Prof. Hung-Yu Kao’s research team is indispensable to the establishment of biological knowledge base, and they hope to develop a comprehensive and stable system which can automatically extract gene numbers and promote its use to biologists in the world.
The task of the BioCreative challenge is to use computer system to automatically extract gene numbers from the texts of a large amount of biomedical papers. It is the most difficult task of the recent biological text mining field.
The main purpose of text mining is to automatically, effectively and accurately find the interacting relationship that biologists are interested in. Before locating the relationship, how to find the gene numbers and gene name normalization corresponding to the genetic database is an important issue.
The problem with gene normalization is to mark the genes mentioned in the biological texts and further transform it to standard gene database numbers. The challenge is that there are too many gene aliases without a complete name dictionary and there are confusions of many different species having the same gene name but different gene numbers, thus requiring one to confirm the genes of the species in the texts.
There are 12 participating teams this year, including National Cheng Kung University, Academia Sinica and Yuan Ze University from Taiwan, National Center for Biotechnology Information, Arizona State University, University of Wisconsin, Milwaukee, University of Iowa, and University of Colorado at Denver from United States, Tsinghua University from China, University of Zurich from Germany, Aveiro University from Portugal and University of Tokyo from Japan.
Among them, many are developing teams of renowned text mining systems, such as GENIA, GNAT, Linnaeus, AIIA-GMT, BIOSMILE and BioThesaurus.
The Critical Assessment for Information Extraction in Biology challenge is a community-wide effort for evaluating text mining and information extraction systems applied to the biomedical domain, bringing together stakeholders from the biocuration community with researchers from text mining and natural language processing applied to the biomedical literature.
Assistant Prof. Hung-Yu Kao of Department of Computer Science and Information Engineering at National Cheng Kung University (NCKU), Tainan, Taiwan, and his supervised doctoral student Chih-Hsuan Wei have shined at the 3rd Critical Assessment of Information Extraction in Biology (BioCreative) challenge in September, 2010, in Bethesda, United States, winning first place in the 5 assessments of the total 9 assessments of gene normalization task with their network model AINet.
The Critical Assessment of Information Extraction systems in Biology, a renowned competition in the field of biological text mining, is initiated by the Centro Nacional de Investigaciones Oncologicas (Spanish National Cancer Research Centre), which is led by Dr. Alfonso Valencia of Spanish National Bioinformatics Institute, Spain, in 2003.
The challenge covers a variety of important issues of contemporary biological text mining, and the results and methods of the winning teams will be invited to be published in renowned journals. For example, the results of the previous challenge were published in the 2010 Nature Biotechnology.
Assistant Prof. Hung-Yu Kao and his supervised doctoral student Chih-Hsuan Wei have participated in the task of gene normalization in the challenge and developed a capture mining system, which does not require training data, by using a method of inference network integrating probability and statistics, being the only team among the contestants which has adopted the non-machine learning method.
In text which has weak signal, they have enlarged the genetic identification evidence with an innovative statistic inference method and applied the method to extract the accurate gene numbers from the articles. This is an innovative and effective application of inference model on bioinformatics issues.
The evaluation rules of the challenge have adopted the TAP-K, Threshold Average Precision-K of National Center for Biotechnology Information, U.S., to assess 507 documents, and among them, 50 documents that are most difficult to extract are divided into two groups for assessment.
In the 9 assessments of the 3 groups, each setting 3 different K-values (K=5, K=10 and K=20), Associate Prof. Hung-Yu Kao and doctoral student Chih-Hsuan Wei have put on outstanding performance in the task of gene normalization, winning first place in 5 assessments of the total 9 assessments and also first place in the overall score.
Currently, the large biological knowledge base still relies on professional force extracting knowledge from the texts and the cost is extremely high.
The network model AINet developed by Associate Prof. Hung-Yu Kao’s research team is indispensable to the establishment of biological knowledge base, and they hope to develop a comprehensive and stable system which can automatically extract gene numbers and promote its use to biologists in the world.
The task of the BioCreative challenge is to use computer system to automatically extract gene numbers from the texts of a large amount of biomedical papers. It is the most difficult task of the recent biological text mining field.
The main purpose of text mining is to automatically, effectively and accurately find the interacting relationship that biologists are interested in. Before locating the relationship, how to find the gene numbers and gene name normalization corresponding to the genetic database is an important issue.
The problem with gene normalization is to mark the genes mentioned in the biological texts and further transform it to standard gene database numbers. The challenge is that there are too many gene aliases without a complete name dictionary and there are confusions of many different species having the same gene name but different gene numbers, thus requiring one to confirm the genes of the species in the texts.
There are 12 participating teams this year, including National Cheng Kung University, Academia Sinica and Yuan Ze University from Taiwan, National Center for Biotechnology Information, Arizona State University, University of Wisconsin, Milwaukee, University of Iowa, and University of Colorado at Denver from United States, Tsinghua University from China, University of Zurich from Germany, Aveiro University from Portugal and University of Tokyo from Japan.
Among them, many are developing teams of renowned text mining systems, such as GENIA, GNAT, Linnaeus, AIIA-GMT, BIOSMILE and BioThesaurus.
The Critical Assessment for Information Extraction in Biology challenge is a community-wide effort for evaluating text mining and information extraction systems applied to the biomedical domain, bringing together stakeholders from the biocuration community with researchers from text mining and natural language processing applied to the biomedical literature.
Click Num:
Share

