Authorship Verification based on Syntax Features

Warning

This publication doesn't include Faculty of Arts. It includes Faculty of Informatics. Official publication website can be found on muni.cz.
Authors

RYGL Jan ZEMKOVÁ Kristýna KOVÁŘ Vojtěch

Year of publication 2012
Type Article in Proceedings
Conference Proceedings of the Sixth Workshop on Recent Advances in Slavonic Natural Language Processing, RASLAN 2012
MU Faculty or unit

Faculty of Informatics

Citation
Web
Field Linguistics
Keywords authorship verification;syntactic analysis;SET;machine learning
Description Authorship verification is wildly discussed topic at these days. In the authorship verification problem, we are given examples of the writing of an author and are asked to determine if given texts were or were not written by this author. In this paper we present an algorithm using syntactic analysis system SET for verifying authorship of the documents. We propose three variants of two-class machine learning approach to authorship verification. Syntactic features are used as attributes in suggested algorithms and their performance is compared to established word-lenth distribution features. Results indicate that syntactic features provide enough information to improve accuracy of authorship verification algorithms.
Related projects:

You are running an old browser version. We recommend updating your browser to its latest version.