Ir al contenido principal

Entradas

Publicación en revista: A method for K-Means seeds generation applied to text mining

Publicado en Statistical Methods and Applications In this paper, a methodology is proposed in order to produce a set of seeds later used as a starting point to K-Means-type unsupervised classification algorithms for text mining. Our proposal involves using the eigenvectors obtained from principal component analysis to extract initial seeds, upon appropriate treatment for search of lightly overlapping clusters which are also clearly identified by keywords. This work is motivated by the interest of the authors in the problem of identification of topics and themes previously unknown in short texts. Therefore, in order to validate the goodness of this method, it was applied on a sample of labeled e-mails (NG20) representing a gold standard within the field of text mining. Specifically, some corpora referenced in the literature have been used, configured in accordance to a mix of topics contained in the sample. The proposed method improves on the results of other state-of-the-art methods to whi...

Concesión del proyecto ATRECSIDE: Algoritmos para el Reconocimiento y la Extraccción de Contenido Semántico en Imágenes de Documentos

Ministerio de Economía y Competitividad Este proyecto se ha centrado en el estudio de la aplicación de las más recientes técnicas de aprendizaje automático para la extracción de contenido semántico en documentos digitalizados que incluyen texto manuscrito y mecanografiado, fotografías e ilustraciones. Dicho objetivo está en línea con las crecientes necesidades actuales de traspaso de información desde el mundo físico al mundo digital. El principal avance derivado del proyecto ha sido el desarrollo de un sistema de reconocimiento de texto manuscrito offline que es estado del arte al compararlo con otros sistemas similares de la bibliografía. Este desarrollo se ha sustentado sobre numerosas publicaciones en revistas y congresos internacionales. En particular se ha realizado 6 publicaciones en revistas JCR, 2 en libros de la serie Lecure Notes y 4 en congresos internacionales. Además, se ha dirigido la digitalización del Archivo de Documentos Históricos de la empresa Osborne, creando ...

Publicación en revista: Beef identification in industrial slaughterhouses using machine vision techniques

Publicado en Spanish Journal of Agriculture Research (SJAR) Accurate individual animal identification provides the producers with useful information to take management decisions about an individual animal or about the complete herd. This identification task is also important to ensure the integrity of the food chain. Consequently, many consumers are turning their attention to issues of quality in animal food production methods. This work describes an implemented solution for individual beef identification, taking in the time from cattle shipment arrival at the slaughterhouse until the animals are slaughtered and cut up. Our beef identification approach is image-based and the pursued goals are the correct automatic extraction and matching between some numeric information extracted from the beef ear-tag and the corresponding one from the Bovine Identification Document (BID). The achieved correct identification results by our method are near 90%, by considering the practica...

Publicación en revista: Off-line handwritten signature detection by analysis of evidence accumulation

Publicado en International Journal on Document Analysis and Recognition (IJDAR) One fundamental step in off-line handwritten signature verification is the detection of the signature position within the document image. This paper introduces an original approach for signature position detection. The method is based on an accumulative evidence technique, searching the region that maximizes some measure of correspondence with a given reference signature. This measure is based on the similarity of the slope marked out by each of the strokes in the signature. Experiments have shown that the method can be used on real documents, such as bank checks, where images have a high noise level due to background interferences (i.e. machine or handwritten texts, stamps, and lines). The proposed method is robust to variability in the size of the signatures and has the advantage of using only one reference signature per person.

Publicación en revista: Fuzzy shape-memory snakes for the automatic off-line signature verification problem

Publicado en Fuzzy Sets and Systems This paper introduces an adapted fuzzy snake approach for efficiently solving some of the practical constraints in the off-line signature verification problem. Our method is called fuzzy shape-memory snakes due to its resemblance to shape-memory alloys, which are metals that in high-temperature conditions can remember their original shape. In our approach, the snake also “remembers” its geometry during its iterative adjustment to a test signature. Off-line signature verification aims to establish the degree of genuineness of a given test signature when compared to a reference signature. Due to the shape and size variability in signatures of the same subject, a system with tolerance to imprecision and also with some “memory” of its initial configured shape, would be very useful for this complex verification problem. To our knowledge, snakes and other active contour models have not been previosly applied to the offline signature verifica...

Publicación en congreso: Robust off-line signature verification using compression networks and positional cuttings

Publicado en 2003 IEEE XIII Workshop on Neural Networks for Signal Processing A novel robust technique for the off-line signature verification problem in practical real conditions is presented. The technique is based on the use of compression neural networks, and in the automatic generation of the training set from only one signature for each writer. Our proposal incorporates a new kind of acceptance/rejection rule, which is based on the similarity between subimages or positional cuttings of a test signature and the corresponding representation stored in the class compression network. Experimental results show that the proposed technique reduces significantly the false acceptation rate (FAR).