Automatic Author Gender Identification From Persian Text Using ParseBert and Linguistic Features

Document Type : Original Article

Authors

1 Assistant Professor, Department of Mathematics, Faculty of Statistics, Mathematics and Computer Science, Allameh Tabatabaei University, Tehran, Iran.

2 Department of linguistics, Faculty of Language and humanities, Bu-Ali Sina University, Hamedan, Iran.

Abstract

The sentences that people use during writing contain valuable information that can be used to identify the author's gender. Meanwhile, the use of deep learning algorithms in natural language processing helps identify hidden patterns in the text. In this research, an attempt is made to design a system for the Persian language that identifies the author's gender by fine-tuning the parameters of the ParsBERT model. For this purpose, first, a corpus of 5,000 documents labeled with gender tags is prepared, and then the author's gender identification system is designed and evaluated using 10-fold cross-validation. Experimental results show that the F-measure of the gender identification task is 76.5%. The proposed method is also compared with classic machine learning methods. It also achieves better results compared with the LSTM model. The results obtained from comparing the corpus prepared in this research with the corpus prepared in previous research for gender identification show an improvement in the system's performance. Thus, the use of new deep learning methods, such as the ParsBERT model, and appropriate data are among the main achievements of this research.
Introduction
Automatic gender classification has become a pivotal research area within computational linguistics due to its wide-ranging applications, including the enhancement of user-friendly digital environments and the improvement of security in online interactions. Since Koppel et al. (2002) pioneered automatic gender identification, technological advancements have increased interest in developing systems that infer demographic attributes, such as gender and age, as well as affective characteristics such as sentiment, from text. For Persian, a language with limited NLP resources, developing such systems is crucial for creating safer and more interactive online spaces, especially given that users can conceal their identities on social media platforms. This study stands out by introducing a gender-labeled emotional corpus and employing ParsBERT, a transformer-based model specifically designed for Persian, to capture subtle linguistic patterns that distinguish between male and female writing styles. The research examines whether ParsBERT can accurately identify the gender of authors in Persian texts and whether the use of an emotionally labeled corpus can enhance classification performance compared with other datasets. The hypotheses of this study are that linguistic features, such as emotional adjectives and adverbs, differ significantly between genders and that deep learning models, such as ParsBERT, outperform traditional machine learning methods.
Theoretical Background
The study is grounded in several decades of research on gender differences in language use. Lakoff (1975) argued that women tend to use specific linguistic patterns, such as confirmatory phrases (e.g., “isn’t it?”), which may reflect social dynamics. Tannen (1994) expanded on this perspective by identifying collaborative speech styles among women, emphasizing empathy, in contrast to competitive styles among men, which focus more on information transfer. Holmes (2013) highlighted the influence of social, cultural, and situational factors, cautioning against overgeneralizing gender differences. Brown and Levinson (1987) associated linguistic politeness strategies with social power dynamics. Recent studies, such as Brody (2013), have explored multifaceted factors, including age and ethnicity, while Chai et al. (2016) have examined online communication. Women have been reported to use more descriptive and emotional vocabulary, such as adjectives (“beautiful,” “amazing”) and intensifiers (“very,” “truly”), while men have been reported to favor functional, direct language (Newman et al., 2008). These differences provide a basis for feature selection in automated gender identification, while acknowledging the complexity of syntactic, semantic, and contextual linguistic patterns.
Related Work
Previous research on gender identification has covered various languages and methodological approaches. Cheng et al. (2011) achieved 85.1% accuracy in English using statistical features and classifiers, including Naive Bayes, Support Vector Machines (SVM), and Decision Trees. In Arabic, Alsmearat et al. (2017) employed bag-of-words representations and morphological features, while Rangel et al. (2019) found that SVM and neural networks were effective. Sboev et al. (2018) used convolutional and LSTM networks for Russian texts, which outperformed traditional methods. In Persian, Moradi and Bahrani (2015) reported 73.8% accuracy using psycholinguistic and stylistic features with traditional classifiers. Sajedi and Taslimi (2019) achieved 89.5% accuracy using Bayesian Random Forests on blog data. Zarifi and Naghavi (2022) reported an accuracy of 81.09% using conceptual vectorization, while Hekmatianzadehpour and Jalali Bidgoli (2019) obtained 85.7% accuracy by incorporating punctuation and n-grams. These studies highlight the potential of combining linguistic features with machine learning approaches, but they also underscore the relatively limited attention given to Persian-specific systems, which is the gap addressed by this research.
Methodology
The methodology focuses on constructing a gender-labeled emotional corpus and applying ParsBERT in conjunction with traditional classifiers. The corpus, adapted from Sadeghi et al. (2021), includes 23,000 Persian sentences categorized into five emotion categories (anger, sadness, joy, fear, and surprise), of which 5,000 documents (2,500 for each gender) were manually labeled according to gender. Preprocessing involves normalization, stemming, and stop-word removal to address Persian-specific challenges, such as inconsistent spacing and orthographic variations (Ghayoomi et al., 2009). For text representation, Word2Vec and bag-of-words approaches are employed, with 100-dimensional vectors representing the average word embeddings of each document. Classification models include Naive Bayes, SVM, Decision Trees (J48), LSTM, and ParsBERT. ParsBERT, a BERT-based model developed for Persian, is fine-tuned with a batch size of 32, a learning rate of 10⁻⁴, and a feed-forward layer with softmax activation for binary classification (female vs. male). It employs 12 hidden layers, 12 attention heads, and a 768-dimensional hidden representation, and is optimized using the Adam optimizer (Kingma & Ba, 2014). Performance is evaluated using 10-fold cross-validation with precision, recall, and F1-score as evaluation metrics.
Evaluation
ParsBERT achieved the highest performance, with an accuracy of 77%, a recall of 76%, and an F1-score of 76.5%, surpassing traditional classifiers and the LSTM model. Its superior performance can be attributed to its transformer-based architecture, which is capable of capturing deep semantic relationships and subtle stylistic differences, such as pronoun usage, sentence length, and emotional tone, that are not readily captured by TF-IDF or Word2Vec embeddings. A comparison with the corpus developed by Moradi and Bahrani (2015), which comprises conversational texts, reviews, and narratives, indicates that the model performs better on the emotional corpus. This improvement may be attributed to the corpus’s emphasis on emotionally charged texts, which may provide more salient gender-related linguistic patterns.
The error analysis revealed three main types of errors: neutral texts that are difficult to distinguish based on gender, texts in which one gender adopts stylistic patterns associated with the other gender, and errors arising from model limitations resulting from the relatively small size of the dataset. Increasing the size and diversity of the corpus may help reduce these errors and consequently improve overall classification performance.
Conclusion
This study contributes to the advancement of automatic gender identification in Persian by introducing a gender-labeled emotional corpus and leveraging the deep-learning capabilities of ParsBERT. The achieved accuracy of 77% demonstrates the effectiveness of transformer-based models in capturing nuanced linguistic differences in comparison with traditional methods. The emotional corpus may further enhance classification performance by providing linguistically rich, emotionally charged texts that highlight gender-specific language patterns. The publicly available corpus provides a valuable resource for future research in Persian NLP. More broadly, this study contributes to the development of NLP resources for low-resource languages and may support more effective human–machine interactions in digital environments.
Future research could focus on developing larger and more diverse corpora and investigating more advanced transformer-based architectures to further improve classification accuracy and generalizability, while addressing the remaining challenges in Persian NLP.
Ethical Considerations
Not applicable
Funding
Not applicable
Conflict of interest
The authors declare no conflict of interest

Keywords

Main Subjects


 
Alsmearat, K., Al-Ayyoub, M., Al-Shalabi, R., & Kanaan, G. G. (2017). Author gender identification from Arabic text. Journal of Information Security and Applications, 35, 85–95.
Alvarez-Carmona, M. A., Pellegrin, L., Montes-y-Gómez, M., Sánchez-Vega, F., Escalante, H. J., López-Monroy, A. P., Villasenor-Pineda, L., & Villatoro-Tello, E. (2018). A visual approach for age and gender identification on Twitter. Journal of Intelligent and Fuzzy Systems, 34(5), 3133–3145.
Brody, L. R. (1993). On understanding gender differences in the expression of emotion: Gender roles, socialization, and language. In S. L. Ablon, D. Brown, E. J. Khantzian, & J. E. Mack (Eds.), Human feelings: Explorations in affect development and meaning (pp. 87–121). Analytic Press.
Brown, P., & Levinson, S. C. (1987). Politeness: Some universals in language usage (Studies in Interactional Sociolinguistics, Vol. 4). Cambridge University Press.
Chai, C., Wu, X., Shen, D., Li, D., & Zhang, K. (2016). Gender differences in the effect of communication on college students’ online decisions. Computers in Human Behavior, 65, 176–188.
Cheng, N., Chandramouli, R., & Subbalakshmi, K. P. (2011). Author gender identification from text. Digital Investigation, 8(1), 78–88.
Devlin, J., Chang, M. W. Lee, K., & Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. https://doi.org/10.48550/arXiv.1810.04805
Farahani, M., Gharachorloo, M., Farahani, M., & Manthouri, M. (2021). ParseBert: Transformer-based model for Persian language understanding. Neural Processing Letters, 53(6), 3831–3847.
Ghayoomi, M., Momtazi, S., & Bijankhan, M. (2009). A study of corpus development for Persian. International Journal on Asian Language Processing, 20(1), 17–33.
Hekmatian Zadehpour, S., & Jalali Bidgoli, A. (2019). Automatic gender recognition of authors of comments written in Persian. Third National Conference on Knowledge and Technology of Electrical, Computer and Mechanical Engineering of Iran. https://civilica.com/doc/925633 (in Persian)
Holmes, J. (2013). Women, men and politeness. Routledge.
Hunter, D., Gambell, T., & Randhawa, B. (2005). Gender gaps in group listening and speaking: Issues in social constructivist approaches to teaching and learning. Educational Review, 57(3), 329–355. https://psycnet.apa.org/doi/10.1080/00131910500149416
Kendall, S., & Tannen, D. (1997). Gender and language in the workplace. In R. Wodak (Ed.), Gender and discourse (pp. 81–105). Sage. https://doi.org/10.4135/9781446250204.n5
Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. https://doi.org/10.48550/arXiv.1412.6980
Koppel, M., Argamon, S., & Shimoni, A. R. (2002). Automatically categorizing written texts by author gender. Literary and Linguistic Computing, 17(4), 401–412.
Lakoff, R. (1975). Language and woman’s place. Language in Society, 2(1), 45–79. https://doi.org/10.1017/S0047404500000051
Moradi, M., & Bahrani, M. (2015). Automatic Gender Identification in Persian Text. Signal and Data Processing, 12(4), 83–94. SID. https://sid.ir/paper/160742/fa (in Persian)
Newman, M. L., Groom, C. J., Handelman, L. D., & Pennebaker, J. W. (2008). Gender differences in language use: An analysis of 14,000 text samples. Discourse Processes, 45(3), 211–236. https://psycnet.apa.org/doi/10.1080/01638530802073712
Rangel, F., Rosso, P., Charfi, A., Zaghouani, W., Ghanem, B., & Sánchez-Junquera, J. (2019). Overview of the track on author profiling and deception detection in Arabic. CEUR Workshop Proceedings, 2517, 70–83.
Sadeghi, S. S., Khotanlou, H., & Rasekh Mahand, M. (2021). Automatic Persian text emotion detection using cognitive linguistic and deep learning. Journal of AI and Data Mining, 9(2), 169–179. https://doi.org/10.22044/jadm.2020.9992.2136
Safara, F., Mohammed, A. S., Potrus, M., Ali, S., Quan, T. T., Souri, A., Janenia, F., & Hosseinzadeh, M. (2020). An author gender detection method using whale optimization algorithm and artificial neural network. IEEE Access, 8, 48428–48437. https://doi.org/10.1109/ACCESS.2020.2973509
Sajedi, H., & Taslimi, M. (2019). Author gender identification from text using bayesian random forest. Signal and Data Processing, 16(1), 143–156. https://doi.org/10.29252/jsdp.16.1.143 (in Persian)
Sboev, A., Moloshnikov, l., Gudovskikh, D., Selivanov, A., Rybka, R., & Litvinova, T. (2018). Automatic gender identification of author of Russian text by machine learning and neural net algorithms in case of gender deception. Procedia Computer Science, 123, 417–423. https://doi.org/10.1016/j.procs.2018.01.064
Simaki, V., Mporas, l., & Megalooikonomou, V. (2016). Evaluation and sociolinguistic analysis of text features for gender and age identification. American Journal of Engineering and Applied Sciences, 9(4), 868-876. https://doi.org/10.3844/ajeassp.2016.868.876
Tannen, D. (1999). Women and men in conversation. The workings of language: From prescriptions to perspectives (pp. 211–216), Bloomsbury Publishing.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.     https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
Zarifi, A., & Naghavi, M. (2022). Gender identification of short text author using conceptual vectorization. Multimedia Tools & Applications, 82(11), 1–17.