Timur Sattarov
4 results
Now showing 1 - 4 of 4
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, FinDiff: Diffusion Models for Financial Tabular Data Generation(Association for Computing Machinery (ACM), 2023); ; ; The sharing of microdata, such as fund holdings and derivative instruments, by regulatory institutions presents a unique challenge due to strict data confidentiality and privacy regulations. These challenges often hinder the ability of both academics and practitioners to conduct collaborative research effectively. The emergence of generative models, particularly diffusion models, capable of synthesizing data mimicking the underlying distributions of real-world data presents a compelling solution. This work introduces 'FinDiff', a diffusion model designed to generate real-world financial tabular data for a variety of regulatory downstream tasks, for example economic scenario modeling, stress tests, and fraud detection. The model uses embedding encodings to model mixed modality financial data, comprising both categorical and numeric attributes. The performance of FinDiff in generating synthetic tabular financial data is evaluated against state-of-the-art baseline models using three real-world financial datasets (including two publicly available datasets and one proprietary dataset). Empirical results demonstrate that FinDiff excels in generating synthetic tabular financial data with high fidelity, privacy, and utility.Type:conference paperJournal:4th ACM International Conference on AI in FinanceScopus© Citations 55 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Federated and Privacy-Preserving Learning of Accounting Data in Financial Statement Audits(Association for Computing Machinery (ACM), 2022-11-01); ; The ongoing 'digital transformation' fundamentally changes audit evidence's nature, recording, and volume. Nowadays, the International Standards on Auditing (ISA) requires auditors to examine vast volumes of a financial statement's underlying digital accounting records. As a result, audit firms also `digitize' their analytical capabilities and invest in Deep Learning (DL), a successful sub-discipline of Machine Learning. The application of DL offers the ability to learn specialized audit models from data of multiple clients, e.g., organizations operating in the same industry or jurisdiction. In general, regulations require auditors to adhere to strict data confidentiality measures. At the same time, recent intriguing discoveries showed that large-scale DL models are vulnerable to leaking sensitive training data information. Today, it often remains unclear how audit firms can apply DL models while complying with data protection regulations. In this work, we propose a Federated Learning framework to train DL models on auditing relevant accounting data of multiple clients. The framework encompasses Differential Privacy and Split Learning capabilities to mitigate data confidentiality risks at model inference. Our results provide empirical evidence that auditors can benefit from DL models that accumulate knowledge from multiple sources of proprietary client data.Type:conference paperScopus© Citations 21 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Multi-view Contrastive Self-Supervised Learning of Accounting Data Representations for Downstream Audit Tasks(Association for Computing Machinery (ACM), 2021-11-03); ; International audit standards require the direct assessment of a financial statement’s underlying accounting transactions, referred to as journal entries. Recently, driven by the advances in artificial intelligence, deep learning inspired audit techniques have emerged in the field of auditing vast quantities of journal entry data. Nowadays, the majority of such methods rely on a set of specialized models, each trained for a particular audit task. At the same time, when conducting a financial statement audit, audit teams are confronted with (i) challenging time-budget constraints, (ii) extensive documentation obligations, and (iii) strict model interpretability requirements. As a result, auditors prefer to harness only a single preferably ‘multi-purpose’ model throughout an audit engagement. We propose a contrastive self-supervised learning framework designed to learn audit task invariant accounting data representations to meet this requirement. The framework encompasses deliberate interacting data augmentation policies that utilize the attribute characteristics of journal entry data. We evaluate the framework on two real-world datasets of city payments and transfer the learned representations to three downstream audit tasks: anomaly detection, audit sampling, and audit documentation. Our experimental results provide empirical evidence that the proposed framework offers the ability to increase the efficiency of audits by learning rich and interpretable ‘multi-task’ representations.Type:conference paperScopus© Citations 10 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Learning Sampling in Financial Statement Audits using Vector Quantised Variational Autoencoder Neural Networks(Association of Computing Machinery (ACM), 2020-10-01); ; ; ;Reimer, BerndThe audit of financial statements is designed to collect reasonable assurance that an issued statement is free from material misstatement ('true and fair presentation'). International audit standards require the assessment of a statements' underlying accounting relevant transactions referred to as 'journal entries' to detect potential misstatements. To efficiently audit the increasing quantities of such journal entries, auditors regularly conduct an 'audit sampling' i.e. a sample-based assessment of a subset of these journal entries. However, the task of audit sampling is often conducted early in the overall audit process, where the auditor might not be aware of all generative factors and their dynamics that resulted in the journal entries in-scope of the audit. To overcome this challenge, we propose the use of a Vector Quantised-Variational Autoencoder (VQ-VAE) neural networks to learn a representation of journal entries able to provide a comprehensive 'audit sampling' to the auditor. We demonstrate, based on two real-world city payment datasets, that such artificial neural networks are capable of learning a quantised representation of accounting data. We show that the learned quantisation uncovers (i) the latent factors of variation and (ii) can be utilised as a highly representative audit sample in financial statement audits.Type:conference paperJournal:Proceedings of the First ACM International Conference on AI in FinanceScopus© Citations 8