<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="https://nkundiushuti.github.io/assets/xslt/rss.xslt" ?>
<?xml-stylesheet type="text/css" href="https://nkundiushuti.github.io/assets/css/rss.css" ?>
<rss version="2.0" xmlns:atom="https://www.w3.org/2005/Atom">
	<channel>
		<title>Marius Miron</title>
		<description>I am a Senior AI Research Scientist at Earth Species Project.</description>
		<link>https://nkundiushuti.github.io/</link>
		<atom:link href="https://nkundiushuti.github.io/feed.xml" rel="self" type="application/rss+xml" />
		
			<item>
				<title>Masato</title>
				<link>https://nkundiushuti.github.io/blog/masato/</link>
				<pubDate>Wed, 15 Apr 2026 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;Last night I received the sad news of my colleague and friend Masato Hagiwara passing. I was considering writing something soon after I read his last &lt;a href=&quot;https://masatohagiwara.net/lcl.html&quot;&gt;post&lt;/a&gt; about language, curiosity, and life.&lt;/p&gt;

&lt;p&gt;Masato was a deeply caring, brilliant, multi-faceted person. He was a scientist but also a musician, to some extent a linguist, always challenging the boundaries of human knowledge. There are certainly many things I don’t know about him, so if I were to choose a single word to define him, that would be “curiosity”. Even in his last days and throughout the battle with cancer, he had this child-like curiosity. There was something very gentle yet nevertheless energizing and motivating about it (it kinda reminded me of Carl Sagan in that aspect).&lt;/p&gt;

&lt;p&gt;We intersected a lot in his passion for science. I decided to join ESP just after he and Jen-Yu interviewed me. I absolutely loved the mindset of pushing the boundaries of science, and the environment was effervescent; every meeting was sparkling with new ideas, new experiments. OMG! They code so fast! Hell yeah! I want to work with these people! We also intersected a lot regarding music. Being completely remote, we met a few times, and when it was possible, we played music. He was an excellent pianist and inclusive with less skilled musicians like me. I still play some jazz standards I was practicing before jamming with him. I will always remember Masato when playing them. I didn’t know he enjoyed a diversity of music styles from many parts of the world, and I was so happy to read that in his last post.&lt;/p&gt;

&lt;p&gt;He was diagnosed 3 years ago, just after I joined ESP, but he decided to continue his normal life. I remember the shock of finding out about it; in my country, that diagnosis is a death sentence. Things got insanely more complicated after his daughter was diagnosed with a serious condition, fought hard, and was finally cured last year. As a parent, I can’t imagine anything more devastating. He and his family decided to approach it differently, and this can be summarized in Japanese by the word yomei, to keep living. And this was a life lesson for me. Indeed, any news can be brutal, and people often decide with this emotional charge. You &lt;a href=&quot;https://www.goodreads.com/book/show/96884.The_Happiness_Hypothesis&quot;&gt;learn to live your new life&lt;/a&gt;, and things get more stable after a while. You approach your new life with a clear mind.&lt;/p&gt;

&lt;p&gt;Just after the diagnosis, he worked on the next generation of the &lt;a href=&quot;https://arxiv.org/abs/2210.14493&quot;&gt;AVES&lt;/a&gt; model, Bird-AVES, which inspired and drove our latest effort &lt;a href=&quot;https://github.com/earthspecies/avex&quot;&gt;AVEX&lt;/a&gt;. He probably did &lt;a href=&quot;https://arxiv.org/abs/2402.03269&quot;&gt;ISPA&lt;/a&gt; as a thought exercise and finished it in a week in the pre-LLM era. He was super involved in larger projects like &lt;a href=&quot;https://arxiv.org/abs/2411.07186&quot;&gt;NatureLM-audio&lt;/a&gt;. In the past months and even weeks, he was working on a benchmark of animal knowledge expertise in LLMs. His influence on where the field is going is hard to quantify, but I think it’s underestimated at the moment. He did all this work because it was important to him, and he wouldn’t have done anything else instead.&lt;/p&gt;

&lt;p&gt;Masato and his family decided to share their journey on the &lt;a href=&quot;https://www.caringbridge.org/site/b0895f78-5320-39d6-96ad-8be3f96e3789&quot;&gt;CaringBridge&lt;/a&gt; website. Many people do not share these life events (I am probably one of them, but more on that in another blog post). Retrospectively, I think the idea was excellent; reading about his day-to-day life helps one accept the current state of things. He liked to hang out on Zoom calls, even for a coffee, and exchange ideas. However, particularly within the last year, I didn’t dare write to him because the condition started to get more serious, and as a work colleague I thought that might be intrusive, but it was somehow reassuring reading about him on CaringBridge. I could feel his presence, although we didn’t speak as much. I hope we will never abandon his memory and legacy and the few lessons about life that he taught us.&lt;/p&gt;

&lt;p&gt;In the past hours, I have been trying to find ways of coping with the gap that he leaves in our lives. It’s like facing the collapse of &lt;a href=&quot;https://www.imdb.com/title/tt12908150/&quot;&gt;a universe of languages, music, facts, hopes, dreams&lt;/a&gt;. This blog post is a step towards that, following one of the many lessons I learned from Masato: being open. I am not a religious person. I can just hope to meet him in the future in a memory-less form of particles colliding in a never-ending universe where all stars are pulled apart more and more. But if the Masato we want to meet is represented by his memory, he is always with us; we can meet him every day by remembering him. So lucky to have met you, Masato!&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/masato/</guid>
			</item>
		
			<item>
				<title>Biodenoising</title>
				<link>https://nkundiushuti.github.io/blog/biodenoising/</link>
				<pubDate>Tue, 01 Oct 2024 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/biodenoising.jpg&quot; alt=&quot;biodenoising&quot; /&gt;&lt;/p&gt;

&lt;p&gt;We published the arxiv preprint of our paper on denoising animal vocalizations without clean data.&lt;/p&gt;

&lt;div class=&quot;alert-box text &quot;&gt;&lt;p&gt;Marius Miron, Sara Keen, Jen-Yu Liu, Benjamin Hoffman, Masato Hagiwara, Olivier Pietquin, Felix Effenberger, Maddie Cusimano, “Biodenoising: animal vocalization denoising without access to clean data”&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;Speech enhancement e.g. denoising speech signals, currently requires large amounts of isolated speech. This is not the case for animal vocalizations. Data recorded in the wild is often noisy, and even biologgers capture self-generated and wind noise. In the lab, self-noise or fan noise are often present.&lt;/p&gt;

&lt;p&gt;How can we denoise animal vocalizations without clean data? The current research relies on estimating the noisy targets in noisier mixture i.e. creating mixtures by adding more noise on top of these targets. However, speech enhancement models already know a lot about patterns in audio time series. So why not use these models to obtain pseudo-clean targets instead of using the noisy targets.&lt;/p&gt;

&lt;p&gt;The paper has the following contributions:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;We introduce a new benchmarking dataset comprising clean vocalizations from different taxa and noise samples.&lt;/li&gt;
  &lt;li&gt;We introduce a separate training set comprising noisy vocalization from existing bioacoustic datasets and noise samples from environmental sounds datasets. The training data can be downloaded and processed using a Python library &lt;strong&gt;&lt;em&gt;biodenoising-datasets&lt;/em&gt;&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;We propose a methodology to leverage speech enhancement models to denoise animal vocalizations without clean data. The experiments can be reproduced using the Python library &lt;strong&gt;&lt;em&gt;biodenoising&lt;/em&gt;&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One important thing, in our experiments we discovered that time scaling e.g. slowing down or speeding up the vocalization can help. Why does this transformation makes sense? We know that with some exceptions the pitch of a vocalization is correlated with the size of the its body. Check out the video below to see how this works for whales and birds.&lt;/p&gt;

&lt;video width=&quot;640&quot; height=&quot;360&quot; id=&quot;player1&quot; preload=&quot;none&quot;&gt;
  &lt;source type=&quot;video/youtube&quot; src=&quot;http://www.youtube.com/watch?v=M5OCCuCIMbA&quot; /&gt;
&lt;/video&gt;

&lt;p&gt;If we can bring up the pitch within the human speech range, we could re-use speech enhancement models. Speech enhancement models learn signal priors such as sparsity, structure of vowels and phonemes that may be useful when denoising animal vocalizations.&lt;/p&gt;

&lt;h2 id=&quot;bibtex&quot;&gt;Bibtex&lt;/h2&gt;

&lt;div class=&quot;alert-box text &quot;&gt;
&lt;p&gt;@misc{miron2024biodenoisinganimalvocalizationdenoising,
      title={Biodenoising: animal vocalization denoising without access to clean data}, 
      author={Marius Miron and Sara Keen and Jen-Yu Liu and Benjamin Hoffman and Masato Hagiwara and Olivier Pietquin and Felix Effenberger and Maddie Cusimano},
      year={2024},
      eprint={2410.03427},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2410.03427}, 
}
“&lt;/p&gt;
&lt;/div&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/biodenoising/</guid>
			</item>
		
			<item>
				<title>Interspeech2024</title>
				<link>https://nkundiushuti.github.io/interspeech2024/</link>
				<pubDate>Thu, 12 Sep 2024 00:00:00 +0000</pubDate>
				<description>&lt;hr /&gt;
&lt;p&gt;layout: page
subheadline: “”
sidebar: right
title: “Interspeech 2024”
teaser: “”
header:
    image_fullwidth: “whale.png”
categories:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;blog
tags:&lt;/li&gt;
  &lt;li&gt;signal processing, speech enhancement, bioacoustics&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This was my first Interspeech in person, during the post-pandemic one in Seul I decided not to travel and chaired a session remotely. This is a large conference with multiple parallel sessions, lots of attendees.&lt;/p&gt;

&lt;p&gt;This is a list of interesting papers that I have seen at Interspeech 2024 in Kos, Greece. Note that this is not a comprehensive list nor is a ranking of the best paper in the programme. The schedule was quite packed so I had to skip many sessions that overlapped to prioritize topics that I currently work on, i.e. bioacoustics.&lt;/p&gt;

&lt;p&gt;To sum up there is a lot of interest in self-supervised models with another benchmark to evaluate these models on, SUPERB 2.0. There are countless papers using these models or even powerful supervised models (Whisper) on various downstream tasks.&lt;/p&gt;

&lt;p&gt;There is also an increasing interest in using neural codecs/tokens rather than spectrograms, waveforms, embeddings from other models. This is reflected in a new challenge The Interspeech 2024 Challenge on Speech Processing Using Discrete Units and several very good papers on this topic.&lt;/p&gt;

&lt;p&gt;Although bioacoustics was a special theme this year I have not seen many papers on this topic. This was surprising but I guess it was good for the VIHAR workshop ESP organized as a satellite event of Interspeech.&lt;/p&gt;

&lt;h4 id=&quot;detection-and-classification-of-bioacoustics-signals&quot;&gt;Detection and classification of bioacoustics signals&lt;/h4&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.09167&quot;&gt;Vision Transformer Segmentation for Visual Bird Sound Denoising&lt;/a&gt; -&amp;gt; laboriously annotated dataset of spectrogram masks that allows to train a vision transformer for denoising&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.08517&quot;&gt;DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species Recognition&lt;/a&gt; -&amp;gt; there was a discussion at the end on what are dialects and whether the geographical stratification of this dataset represents dialects accurately&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/cauzinille24_interspeech.html&quot;&gt;Investigating self-supervised speech models’ ability to classify animal vocalizations: The case of gibbon’s vocal signatures&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/barnhill24_interspeech.html&quot;&gt;ANIMAL-CLEAN – A Deep Denoising Toolkit for Animal-Independent Signal Enhancement&lt;/a&gt; -&amp;gt; train a noisy2noisy model for each dataset, an extension of orca clean on other datasets&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/sheikh24_interspeech.pdf&quot;&gt;Bird Whisperer: Leveraging Large Pre-trained Acoustic Model for Bird Call Classification&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id=&quot;generative-speech-enhancement&quot;&gt;Generative speech enhancement&lt;/h4&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.12194&quot;&gt;Universal Score-based Speech Enhancement with High Content Preservation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id=&quot;speech-benchmarks&quot;&gt;Speech benchmarks:&lt;/h4&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/shi24g_interspeech.html&quot;&gt;ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets&lt;/a&gt; -&amp;gt; included new metrics such as macro-average and standard deviation; added different fine-tuning options; compared against supervised methods and they seem to work worse than self-supervised on the downstream tasks&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id=&quot;neural-codecs&quot;&gt;Neural codecs:&lt;/h4&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/chang24b_interspeech.html&quot;&gt;The Interspeech 2024 Challenge on Speech Processing Using Discrete Units&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/mousavi24_interspeech.html&quot;&gt;How Should We Extract Discrete Audio Tokens from Self-Supervised Models?&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.12434&quot;&gt;Towards Audio Codec-based Speech Separation&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/li24qa_interspeech.html&quot;&gt;On the Effectiveness of Acoustic BPE in Decoder-Only TTS&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/dekel24_interspeech.html&quot;&gt;Exploring the Benefits of Tokenization of Discrete Acoustic Units&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.11037&quot;&gt;NAST: Noise Aware Speech Tokenization for Speech Language Models&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id=&quot;foundation-models&quot;&gt;Foundation models&lt;/h4&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2311.02248&quot;&gt;COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2310.12477&quot;&gt;Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.08402&quot;&gt;Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.isca-archive.org/interspeech_2024/deshmukh24b_interspeech.html&quot;&gt;PAM: Prompting Audio-Language Models for Audio Quality Assessment&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2406.08619&quot;&gt;Self-Supervised Speech Representations are More Phonetic than Semantic&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/interspeech2024/</guid>
			</item>
		
			<item>
				<title>Evaluating causes of algorithmic bias in juvenile criminal recidivism</title>
				<link>https://nkundiushuti.github.io/blog/causes-of-unfairness/</link>
				<pubDate>Wed, 10 Jun 2020 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;We have published an extended version of our ICAIL19 awarded paper in the Journal of AI and Law &lt;a href=&quot;https://link.springer.com/article/10.1007/s10506-020-09268-y/&quot;&gt;(paper)&lt;/a&gt;. This is an interdisciplinary effort involving computer science (machine learning) and social sciences (decision making/psychology):&lt;/p&gt;

&lt;div class=&quot;alert-box text &quot;&gt;&lt;p&gt;Marius Miron, Songül Tolan, Emilia Gomez, Carlos Castillo, “Evaluating causes of algorithmic bias in juvenile criminal recidivism”&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;We study the implication of using machine learning to predict reoffense risk of defendants in prison and we compare that with SAVRY, a structured professional risk assessment tool. In particular, we look at the causes of discrimination: differences in the prevalence of reoffense between protected groups (women vs men, foreigners vs non-foreigners), input/training features, and the machine learning model (neural network, logistic regression, svm etc.).&lt;/p&gt;

&lt;p&gt;Note that in the &lt;a href=&quot;https://dl.acm.org/doi/pdf/10.1145/3322640.3326705&quot;&gt;ICAIL19 paper&lt;/a&gt; we show that ML methods can achieve better predictive performance than SAVRY but at the expense of unfairer outcomes. Furthermore, in comparison to ML tools which give a score, often with no explanation, SAVRY is not actuarial. It is completed by an expert and it is used to help the defendants improve on the problematic areas detected by a simple sum of scores.&lt;/p&gt;

&lt;p&gt;With regards to the feature sets, we have a blog post from last year in which we explain the difference between SAVRY and Non-SAVRY features. Basically, SAVRY features are more related to things that the defendant can change, while Non-SAVRY features are related to demographics and personal history. In the ICAIL19 paper we show that using Non-SAVRY features yields better predictive performance and leads to more unfair outcomes than using SAVRY features.&lt;/p&gt;

&lt;p&gt;For this paper, as well as in ICAIL19, we use the Catalonian data on juvenile recidivism and SAVRY. You can find the link to these data in the paper.&lt;/p&gt;

&lt;p&gt;We equalize base rates (prevalence) between a reference group (non-foreigners or men) and the other groups (foreigners or women). There are a few ways to do this. We choose to oversample the non-reference groups by repeating in the dataset more people who have recidivated or not in order to match the prevalence of the reference group (EBR). This is a simple way of transforming the data and we compare it with a pre-processing mitigation: LFR (Zemel et al. 2013). LFR is based on an optimization procedure which transforms the data with respect to a protected feature (e.g. race) by removing any information about membership with respect to the protected group.&lt;/p&gt;

&lt;p&gt;There are a few ways to measure algorithmic discrimination. In the paper we look at false positive rate disparity (FPRD) and false negative rate disparity (FNRD). For example, a FPRD for foreigners of 2 means that foreigners are twice as likely to be wrongfully classified as recidivists. Similarly, a FPRD for foreigners of 0.5 means that non-foreigners are twice as likely to be wrongfully classified as non-recidivists.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/foreigner.png&quot; alt=&quot;foreigner&quot; /&gt;&lt;/p&gt;

&lt;p&gt;In the plots we compare EBR, LFR and the baseline (e.g. not doing any transformation on the data) for foreigners with respect to non-foreigners using the feature sets SAVRY, Non-SAVRY and All (their combination).&lt;/p&gt;

&lt;p&gt;EBR, ensuring that the base rates are equal between foreigners and non-foreigners in the training data (both groups have the same prevalence of recidivism) reduces this disparity at a similar level with LFR. Moreover, EBR and LFR are effective when using features correlated with recidivism, like the demographic and personal history features present in the Non-SAVRY feature set. Furthermore, we see that the baseline often yields disparity which is not observed for the SAVRY Sum (the simple sum of SAVRY scores) and the Expert evaluations.&lt;/p&gt;

&lt;p&gt;What is the impact of equalizing base rates on the area under the curve (AUC)? We found something interesting here: when effective in terms fairness, EBR and LFR experience a drop in terms of AUC: 0.01 for EBR and 0.06 for LFR. For the SAVRY feature set, EBR and LFR are not effective in reducing disparity and their AUC does not drop. In this case, there is a clear trade-off between predictive performance and fairness.&lt;/p&gt;

&lt;p&gt;What if we use EBR? It’s simple, it reduces disparity, it doesn’t lose so much in terms of AUC. Well, first of all we use EBR with respect to a protected feature, e.g. foreigner. There is nothing ensuring that it will be fair with respect to sex, national groups, or any other protected feature that we care about. Moreover, by doing mitigation (LFR and EBR) you may end up discriminating even more with respect to other groups or even with under-represented sub-groups (e.g. you may discriminate against the Maghrebi subgroup of the foreigners group)!&lt;/p&gt;

&lt;p&gt;Second, using LIME we looked at which features are important for EBR, LFR, and the baseline. With the exception of the SAVRY feature set, ML relies on demographic and personal history features. While doing EBR does not change much, the top 10 important features LFR are mostly SAVRY features.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Baseline:
    &lt;ul&gt;
      &lt;li&gt;Foreigner   242.83  13.82&lt;/li&gt;
      &lt;li&gt;Sex*    233.00  22.87&lt;/li&gt;
      &lt;li&gt;National group* 221.75  13.98&lt;/li&gt;
      &lt;li&gt;Province of residence*  192.16  13.21&lt;/li&gt;
      &lt;li&gt;No. of maincrimes*  167.58  9.87&lt;/li&gt;
      &lt;li&gt;Age maincrime*  167.92  11.23&lt;/li&gt;
      &lt;li&gt;Province of execution*  165.53  9.38&lt;/li&gt;
      &lt;li&gt;Maincrime program sentence* 163.08  7.74&lt;/li&gt;
      &lt;li&gt;No. of prior crimes*    167.63  12.34&lt;/li&gt;
      &lt;li&gt;Maincrime program duration*&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;EBR on attribute “nationality”
    &lt;ul&gt;
      &lt;li&gt;Foreigner*  228.52  18.53&lt;/li&gt;
      &lt;li&gt;National group* 213.45  12.75&lt;/li&gt;
      &lt;li&gt;Sex*    218.75  21.10&lt;/li&gt;
      &lt;li&gt;Province of residence*  187.01  13.35&lt;/li&gt;
      &lt;li&gt;No. of prior crimes*    155.59  8.86&lt;/li&gt;
      &lt;li&gt;No. of maincrimes*  155.28  9.76&lt;/li&gt;
      &lt;li&gt;Province of execution*  152.80  8.04&lt;/li&gt;
      &lt;li&gt;Maincrime program*  152.35  8.99&lt;/li&gt;
      &lt;li&gt;Age maincrime*  157.60  15.19&lt;/li&gt;
      &lt;li&gt;Maincrime program duration*&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
  &lt;li&gt;LFR on attribute “nationality”
    &lt;ul&gt;
      &lt;li&gt;Antisocial behaviour    123.50  18.96&lt;/li&gt;
      &lt;li&gt;Attention deficit hyperactivity 123.29  18.85&lt;/li&gt;
      &lt;li&gt;Anger management issues 122.70  18.39&lt;/li&gt;
      &lt;li&gt;Treatment susceptibility    122.58  18.30&lt;/li&gt;
      &lt;li&gt;High interest in school/work    123.38  19.20&lt;/li&gt;
      &lt;li&gt;Pro-social support (by adult)   122.10  18.13&lt;/li&gt;
      &lt;li&gt;Substance abuse 122.50  18.64&lt;/li&gt;
      &lt;li&gt;Personality 122.54  18.73&lt;/li&gt;
      &lt;li&gt;Family dynamics 122.38  18.66&lt;/li&gt;
      &lt;li&gt;Positive/resilience characteristics 121.59  17.95&lt;/li&gt;
    &lt;/ul&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The star symbol denotes Non-SAVRY features which are mostly related to demographic and personal history features.&lt;/p&gt;

&lt;p&gt;Now, rememeber that LFC experienced more AUC drop than EBR. This means that we can’t have good predictive performance and fair outcomes on this problem using machine learning. LFR as our most fair option (in terms of ML) has similar or worse AUC to the simple SAVRY sum of scores or Expert evaluation (human-in-the-loop). Note that in comparison to SAVRY sum, LFR cannot be fair with respect to all protected features (foreigner and sex). Mitigation itself is problematic in this case.&lt;/p&gt;

&lt;p&gt;In conclusion, discrimination in ML for juvenile recidivism prediction can be explained partially by the difference in base rates. In addition, the nature of features used for training matters: SAVRY vs Non-SAVRY. Having features correlated with demographics and personal history may accentuate discrimination. Moreover, there are some differences between the ML algorithms, which point out that the ML algorithms may pick of different things when classifying someone as recidivist/non-recidivist.&lt;/p&gt;

&lt;p&gt;So the question that policy makers ask themselves is why replace SAVRY with ML under these conditions? What are the areas that are of high risk?&lt;/p&gt;

&lt;p&gt;There is a trend in computer science to advocate for neutrality of technology. However, as science in itself is neutral, there is a grey zone in computer science which contains applications which are far from neutral. Take for instance face recognition with all its inherent biases. Recently, companies as Microsoft and Amazon discontinued their work and support on this area. In my opinion, face recognition is just an application of image classification, and the way it has been designed and evaluated had a lot to do and did not care at all about the ethical implications and biases. Which makes me think whether these applications should be designed at all.&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/causes-of-unfairness/</guid>
			</item>
		
			<item>
				<title>Fairness in machine learning - the case of juvenile criminal justice in Catalonia</title>
				<link>https://nkundiushuti.github.io/blog/juvenile-recidivism/</link>
				<pubDate>Sun, 16 Jun 2019 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;In the HUMAINT project we deal, among other things, with the issue of  fairness in machine learning, which plays a big role in  research communities such as FAT-ML (a workshop organized within the ICML conference) and FAT* (the main conference on this topic). Many of the papers in this field are interdisciplinary, authored by researchers from computer science, social sciences, economics or law. Part of the papers focus on developing fair machine learning systems, understanding the causes of unfairness, and, since it’s a new field, framing the problem correctly for different areas/systems.&lt;/p&gt;

&lt;p&gt;We presented our paper on fairness in juvenile criminal justice (the Catalonian SAVRY)at the International Conference on AI and Law, where it has received the best paper award:&lt;/p&gt;

&lt;div class=&quot;alert-box text &quot;&gt;&lt;p&gt;Songül Tolan, Marius Miron, Emilia Gomez, Carlos Castillo, “Why Machine Learning May Lead to Unfairness: Evidence from Risk Assessment for Juvenile Justice in Catalonia”&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;In this paper we analyze a tool and several machine learning models designed to predict reoffense risk of defendants in prison in terms of predictive performance and fairness. However, we do not stop there. We use interpretable machine learning to trace back discrimination by assessing which features were important when the algorithm was correct or wrong with a category of people (here, foreigners).&lt;/p&gt;

&lt;h2 id=&quot;machine-learning-for-decision-making&quot;&gt;Machine learning for decision making&lt;/h2&gt;

&lt;p&gt;When judges decide whether to detain or release defendants awaiting trial, they must consider the defendants flight risk or the likelihood to reoffend. The act of reoffense after being convicted for another crime in the past is also termed recidivism. Increasingly, criminal justice systems use algorithms to support judge decisions with machine predictions of recidivism risk. These algorithms derive their rules from data on past cases, corresponding information on judge decisions and information on recidivism (usually we allow up to two years after the exit from prison). We evaluate the performance of of algorithmic and human decision making on splits of the same data to tell whether the decision was correct or not.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/blog1.png&quot; alt=&quot;decmaking1&quot; /&gt;&lt;/p&gt;

&lt;p&gt;There are various reasons why we would like to have a machine learning model support decision making: judges are human and  humans have biases, as shown in this book. On the other hand, there are reasons why we should be careful when considering the support of machine learning systems: machine learning models inherit human bias (mostly through data). Technology is not value-neutral and it’s created by people who express a set of values in the things they create. Often this has  unintended consequences, such as discrimination.
Each data point in a dataset represents a criminal case which is characterized by a set of features. Machine learning models learn from these features to separate between recidivists and non-recidivists.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/blog2.png&quot; alt=&quot;decmaking2&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;what-features-did-we-use-to-predict-recidivism-in-this-case&quot;&gt;What features did we use to predict recidivism in this case?&lt;/h2&gt;

&lt;p&gt;Age at main crime, sex, nationality, sentence, previous number of crimes, year of crime, whether probation was given or not etc. These are features related to demographics and criminal history and we represent them with red.&lt;/p&gt;

&lt;p&gt;The literature on fair algorithms mainly derives its fairness concepts from a legal context. Generally, a process or decision is considered fair if it does not discriminate against people on the basis of their membership to a protected group, such as sex or race.  In this case we tested for discrimination against foreigners and on the basis of sex. We detect discrimination in a decision making process (indipendent of a human or algorithmic origin) by testing the decision against a “ground truth”(e.g. recividism occured within two years after release date or not). Processes that are overproportionally correct or wrong for a particular protected feature indicate discrimination.
Risk assessment tools
Tools which assist decision makers are nowadays used in criminal justice, medicine, finance etc. One of the well known tools for recidivism prediction is COMPAS, which stirred some controversy a while ago because arguably it discriminates against African-American people.&lt;/p&gt;

&lt;p&gt;SAVRY stands for Structured Assessment of Violent Risk in Youth and it’s used in many countries across the world to assess the risk of violent recidivism for young people. It has been mainly designed to work well for violent crimes and male juveniles. In contrast to COMPAS, SAVRY is a transparent list of features/terms,and the final risk evaluation remains at the discretion of the human expert..&lt;/p&gt;

&lt;h2 id=&quot;what-features-do-we-have-in-savry&quot;&gt;What features do we have in SAVRY?&lt;/h2&gt;
&lt;p&gt;SAVRY is based on a questionnaire which collects information on: early violence, self-harm violence, home violence, poor school achievement, stress and poor coping mechanisms, substance abuse, criminal parent/caregiver.
However, SAVRY is expensive, as it requires the expertise and the time to gather all that data. We wanted to see if machine learning models are better in predicting recidivism and if they exhibit any discrimination, in comparison to SAVRY.
Experiments
In terms of input features we test the following three sets:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;Static features (demographics and criminal history features) - red features&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;SAVRY features - blue features&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Static + SAVRY - red+blue features&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In terms of machine learning methods, we tested logistic regression, multi-layer perceptron, svm, knn, random forest, naive bayes to report predictive performance (area under the curve), but we stayed with the top two performing ones to evaluate fairness (logistic regression and multi-layer perceptron).&lt;/p&gt;

&lt;p&gt;The dataset we used in our experiments comprises 855 juvenile offenders aged 12-17 in Catalonia  with crimes committed between 2002 -2010. The release was in 2010  and the recidivism status was followed up in 2013 and 2015.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/blog3.png&quot; alt=&quot;decmaking3&quot; /&gt;&lt;/p&gt;

&lt;p&gt;In our experiments the SAVRY Sum (0.64) and Expert (0.66) have a lower predictive power (lower AUC) than the ML models.&lt;/p&gt;

&lt;p&gt;However, when looking at the classification errors done with respect to foreigners (false positive rates for foreigners compared to false positive rates for Spanish), ML is generally more unfair.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/blog4.png&quot; alt=&quot;decmaking4&quot; /&gt;&lt;/p&gt;

&lt;p&gt;In fact, the more non-SAVRY features (red) you introduce, the more you increase the disparity between Spaniards and foreigners. Not using non-SAVRY features is still discriminative (first two bars in the first column) which was expected: many state of the art papers report discrimination even in the absence of problematic features.&lt;/p&gt;

&lt;p&gt;The difference between the three sets of features led us to analyze the importance of the features by using machine learning interpretability. With respect to that, while some models are interpretable by definition, for other black-box models we need post-hoc methods which can give an approximate interpretation. One of these frameworks, possibly the most used is LIME.&lt;/p&gt;

&lt;p&gt;For logistic regression the importance is given by the weights learned by the model:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/blog5.png&quot; alt=&quot;decmaking5&quot; /&gt;&lt;/p&gt;

&lt;p&gt;For multi-layer perceptron, LIME yields the following global importance:
&lt;img src=&quot;https://nkundiushuti.github.io/images/blog6.png&quot; alt=&quot;decmaking6&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Note that when you combine non-Savry (red) and SAVRY (blue) features, the model relies more on the non-SAVRY features on demographics and criminal history.&lt;/p&gt;

&lt;p&gt;As we mentioned, not using the red features has lead to discrimination. This made us think that the causes of discrimination do not lie in the features but in the data itself. It is established  in the literature that one of the important sources of discrimination is data itself. A good indicator of this is the difference in prevalence of recidivism also termed “base rates” between Spaniards and foreigners. Basically, if 46% of the foreigners in your dataset are recidivists compared to solely 32% of the Spaniards, then your model will learn that the features correlated to being a foreigner are important. Such a model will yield more false alarms with respect to foreigners.&lt;/p&gt;

&lt;p&gt;Therefore, we conclude that algorithms that predict recidivism in criminal justice should only be used under the awareness of such fairness issues.&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/juvenile-recidivism/</guid>
			</item>
		
			<item>
				<title>Music Bandwidth Expansion</title>
				<link>https://nkundiushuti.github.io/blog/high-frequency/</link>
				<pubDate>Fri, 14 Sep 2018 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;Shortly after finishing writing my PhD thesis I collaborated with &lt;a href=&quot;https://telecom.inesctec.pt/~mdavies&quot;&gt;Matthew Davies&lt;/a&gt; from INESC TEC, Porto on music bandwidth expansion - basically recovering the high frequency part of the spectrum. This has big implications in encoding and transmitting music and audio in general through bandwidth-limited media. You might recall how the voice is compressed before transmitting it through telephone lines, and then reconstructed. And how you waited and listened to a crappy sounding Mozart tune while waiting to talk with the support department. Well, if your phone had a system to better reconstruct that music piece, then it would have sounded better.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://telecom.inesctec.pt/~mdavies/dafx18/&quot;&gt;&lt;img src=&quot;https://telecom.inesctec.pt/~mdavies/dafx18/sounds/overviewfig.png&quot; alt=&quot;bandwidth-expansion&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can check the &lt;a href=&quot;https://telecom.inesctec.pt/~mdavies/dafx18/&quot;&gt;companion page&lt;/a&gt; for music examples and more details on the method.&lt;/p&gt;

&lt;div class=&quot;alert-box text &quot;&gt;&lt;p&gt;Abstract:
We present a new approach for audio bandwidth extension for music signals using convolutional neural networks (CNNs). Inspired by the concept of inpainting from the field of image processing, we seek to reconstruct the high-frequency region (i.e., above a cutoff frequency) of a time-frequency representation given the observation of a band-limited version. We then invert this reconstructed time-frequency representation using the phase information from the band-limited input to provide an enhanced musical output. We contrast the performance of two musically adapted CNN architectures which are trained separately using the STFT and the invertible CQT. Through our evaluation, we demonstrate that the CQT, with its logarithmic frequency spacing, provides better reconstruction performance as measured by the signal to distortion ratio.&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;The &lt;a href=&quot;https://telecom.inesctec.pt/~mdavies/pdfs/MironDavies18-dafx.pdf&quot;&gt;paper&lt;/a&gt; was presented in the Digital Audio Effects 2018 conference in Portugal.&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/high-frequency/</guid>
			</item>
		
			<item>
				<title>Interpretability</title>
				<link>https://nkundiushuti.github.io/blog/interpretability/</link>
				<pubDate>Wed, 04 Apr 2018 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;Interpretability is a requisite of machine learning systems which allows for an explanation of their decisions or functionality. The machine learning literature does not have a consensus on a definition of interpretability and state of the arts methods evaluate it through proxy characteristics. Fortunately, this topic benefits from two position papers [Doshi-Velez and Kim (2017), Lipton (2016)] which try to formalize it. I will present the main points of these papers. Then, I will discuss interpretability in medicine as a strictly-regulated field, and the potential for confirmation bias in personalized explanations.&lt;/p&gt;

&lt;h2 id=&quot;position-papers&quot;&gt;Position papers&lt;/h2&gt;

&lt;p&gt;There are two main position papers on the topic of interpretability:&lt;/p&gt;

&lt;p&gt;Doshi-Velez, F., &amp;amp; Kim, B. (2017). Towards A Rigorous Science of Interpretable Machine Learning, (Ml), 1–13. &lt;a href=&quot;https://arxiv.org/abs/1702.08608&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;video width=&quot;640&quot; height=&quot;360&quot; id=&quot;player1&quot; preload=&quot;none&quot;&gt;
  &lt;source type=&quot;video/youtube&quot; src=&quot;https://www.youtube.com/embed/bQfYRcXc9F0&quot; /&gt;
&lt;/video&gt;

&lt;p&gt;Lipton, Z. C. (2016). The Mythos of Model Interpretability, (Whi). &lt;a href=&quot;https://arxiv.org/abs/1606.03490&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Because the literature on interpretability comes with different definitions and ways of evaluating it, both of these position papers underline the need to formalize this term.&lt;/p&gt;

&lt;h2 id=&quot;why-do-we-need-interpretability&quot;&gt;Why do we need interpretability:&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;Scientific understanding -&amp;gt; We need to gain knowledge about a topic. Often we are not solely interested in a method’s accuracy, but also in its informativeness.&lt;/li&gt;
  &lt;li&gt;Safety -&amp;gt; We want to make sure that an AI system is robust to domain mismatch (e.g. a medical AI which works in a different hospital than where it was design which has a different prevalence of a disease), adversarial attacks (through AI generated examples from which a given AI system can learn erroneous decision boundaries, or which can manipulate the decision boundaries of the AI system), errors in training data (we want to detect potential biases in training data)&lt;/li&gt;
  &lt;li&gt;Ethics -&amp;gt; We want to ensure that the algorithms are fair and do not discriminate&lt;/li&gt;
  &lt;li&gt;Mismatch objectives -&amp;gt; ML models usually learn by minimizing a loss function. In many cases it is difficult to translate interpretability in such a function, resulting in an opaque model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Often, interpretability comes as a proxy to various characteristics we seek in a machine learning models. These are the auxiliary or complementary characteristics, which might compete with each other are:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Fairness, un-biasness - we do not want our algorithm to discriminate. For evaluating fairness we have formal criteria; &lt;a href=&quot;https://arxiv.org/abs/1610.02413&quot;&gt;pdf&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Privacy - we want our algorithm to protect the privacy of the data it learns from. This has formal evaluation criteria. &lt;a href=&quot;https://arxiv.org/abs/0907.3754&quot;&gt;pdf&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Reliability - do we trust our algorithm to be robust in any real-life test scenario? Does it generalize well? Is it vulnerable to adversarial attacks (e.g. AI generated examples which manipulate the decision of another AI)&lt;/li&gt;
  &lt;li&gt;Causality - the associations learned reflect true causes rather than spurious correlations.&lt;/li&gt;
  &lt;li&gt;Trust - algorithms which are right for the right reasons. We can correctly predict the decision boundaries of the system and check where it will go wrong.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Unfortunately, reliability, causality and trust do not have a formal criteria, can’t be enforced explicitly through a mathematical equation or included within a cost function.&lt;/p&gt;

&lt;h2 id=&quot;evaluation-taxonomy&quot;&gt;Evaluation taxonomy:&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;Application-grounded evaluation: in this scenario we have real humans, real tasks e.g. medical imaging -&amp;gt; We are interested in asking questions on the quality of explanation: does it better identify errors, new facts, discrimination&lt;/li&gt;
  &lt;li&gt;Human-grounded evaluation: real humans, simplified tasks -&amp;gt; We are interested in studying what kinds of explanations are better&lt;/li&gt;
  &lt;li&gt;Functionally-grounded evaluation: no humans, proxy tasks -&amp;gt; We can only do these experiments after human-grounded of application grounded evaluation, once we have a model of that
Doshi-Velez and Kim (2017) recommend human produced explanations as a baseline. I am not sure whether this results in confirmation bias. Check the last section of this document.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;interpretability-hypotheses&quot;&gt;Interpretability hypotheses:&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;Global vs local: what are the most relevant features for a particular class vs why has this particular example/data point been classified as an instance of this particular class&lt;/li&gt;
  &lt;li&gt;Area, severity of incompleteness: what part of the problem is incomplete and to what extent? E.g. safety in self-driving cars vs sensors that cause the car to drive off the road by 10cm&lt;/li&gt;
  &lt;li&gt;Time constraints/user cost: how long can the user afford to spend to understand the explanation? What is the cost of an explanation? A personal credit system should be more interpretable than musical recommendation system?&lt;/li&gt;
  &lt;li&gt;Nature of the user expertise: sets the level of sophistication of the explanation: decision makers, researchers, expert users, users, safety engineers, data scientists&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;properties-of-interpretable-models&quot;&gt;Properties of interpretable models:&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;Simulatabilty: the model can be contemplated at once e.g. small linear models&lt;/li&gt;
  &lt;li&gt;Decomposability: each parameter and input feature can be interpreted Algorithmic transparency: linear model vs deep learning&lt;/li&gt;
  &lt;li&gt;Post-hoc interpretability: text, image, local explanations, explanation by example&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;existing-methodsframeworks&quot;&gt;Existing methods/frameworks&lt;/h2&gt;
&lt;h3 id=&quot;post-hoc-interpretability-frameworks&quot;&gt;(Post-hoc) Interpretability frameworks&lt;/h3&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/marcotcr/lime&quot;&gt;LIME&lt;/a&gt; -  local+global explanations, evaluating trust, proxy to detect bias, robustness, can evaluate a multitude of models&lt;/p&gt;

&lt;p&gt;Ribeiro, M. T., Singh, S., &amp;amp; Guestrin, C. (2016). “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. &lt;a href=&quot;https://doi.org/10.1145/1235&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;video width=&quot;640&quot; height=&quot;360&quot; id=&quot;player1&quot; preload=&quot;none&quot;&gt;
  &lt;source type=&quot;video/youtube&quot; src=&quot;https://www.youtube.com/embed/KP7-JtFMLo4&quot; /&gt;
&lt;/video&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/ramprs/grad-cam&quot;&gt;GRADCAM&lt;/a&gt; - local explanations, evaluating reliability/trust, can evaluate convolutional neural networks&lt;/p&gt;

&lt;p&gt;Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., &amp;amp; Batra, D. (2017). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. Proceedings of the IEEE International Conference on Computer Vision, 2017–Octob, 618–626. &lt;a href=&quot;https://doi.org/10.1109/ICCV.2017.74&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.heatmapping.org/&quot;&gt;LRP&lt;/a&gt; - local explanations, evaluating relevant features, can evaluate neural networks&lt;/p&gt;

&lt;p&gt;Montavon, G., Samek, W., &amp;amp; Müller, K. R. (2018). Methods for interpreting and understanding deep neural networks. Digital Signal Processing: A Review Journal, 73, 1–15. &lt;a href=&quot;https://doi.org/10.1016/j.dsp.2017.10.011&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/kundajelab/deeplift&quot;&gt;DeepLIFT&lt;/a&gt; - local explanations, counterfactual,  can evaluate neural networks&lt;/p&gt;

&lt;p&gt;Shrikumar, A., Greenside, P., &amp;amp; Kundaje, A. (2017). Learning Important Features Through Propagating Activation Differences. &lt;a href=&quot;https://arxiv.org/abs/1704.02685&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;video width=&quot;640&quot; height=&quot;360&quot; id=&quot;player1&quot; preload=&quot;none&quot;&gt;
  &lt;source type=&quot;video/youtube&quot; src=&quot;https://www.youtube.com/embed/v8cxYjNZAXc&quot; /&gt;
&lt;/video&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/CSAILVision/NetDissect&quot;&gt;Network Dissection&lt;/a&gt; - local explanations, model interpretability, interpretability = alignment with semantic concepts, can evaluate convolutional neural networks&lt;/p&gt;

&lt;p&gt;Bau, D., Zhou, B., Khosla, A., Oliva, A., &amp;amp; Torralba, A. (2017). Network Dissection: Quantifying Interpretability of Deep Visual Representations. &lt;a href=&quot;https://doi.org/10.1109/CVPR.2017.354&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/kohpangwei/influence-release&quot;&gt;Influence functions&lt;/a&gt; - local explanations, counterfactual, can evaluate a multitude of models&lt;/p&gt;

&lt;p&gt;Koh, P. W., &amp;amp; Liang, P. (2017). Understanding Black-box Predictions via Influence Functions. &lt;a href=&quot;https://arxiv.org/abs/1703.04730&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;video width=&quot;640&quot; height=&quot;360&quot; id=&quot;player1&quot; preload=&quot;none&quot;&gt;
  &lt;source type=&quot;video/youtube&quot; src=&quot;https://www.youtube.com/embed/0w9fLX_T6tY&quot; /&gt;
&lt;/video&gt;
&lt;p&gt;s&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/dtak/rrr&quot;&gt;RRR&lt;/a&gt; - local and global explanation, can evaluate gradient-based methods - neural networks&lt;/p&gt;

&lt;p&gt;Ross, A. S., Hughes, M. C., &amp;amp; Doshi-Velez, F. (2017). Right for the right reasons: Training differentiable models by constraining their explanations. IJCAI International Joint Conference on Artificial Intelligence, 2662–2670. &lt;a href=&quot;https://doi.org/10.24963/ijcai.2017/371&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;video width=&quot;640&quot; height=&quot;360&quot; id=&quot;player1&quot; preload=&quot;none&quot;&gt;
  &lt;source type=&quot;video/youtube&quot; src=&quot;https://www.youtube.com/embed/MMxZlr_L6YE&quot; /&gt;
&lt;/video&gt;

&lt;h3 id=&quot;collections-of-methods&quot;&gt;Collections of methods:&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;a href=&quot;https://github.com/marcoancona/DeepExplain&quot;&gt;DeepExplain&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;a href=&quot;https://github.com/LASER-UMASS/Themis&quot;&gt;Themis&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;a href=&quot;https://github.com/jphall663/interpretable_machine_learning_with_python&quot;&gt;IMLP&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;strictly-regulated-fields-medical-ai&quot;&gt;Strictly-regulated fields: medical AI&lt;/h2&gt;

&lt;p&gt;Interpretation is highly valued in strictly-regulated fields such as medicine. Rather than aiming at an end-to-end solution where the machine learning algorithm gives a diagnostic, ideally replacing the doctor, the medicine AI target sub-tasks which can assist doctors in making a more informed decision or help them become more efficient by filtering some true-positives or false negatives. By targeting sub-tasks the medical AI works at a lower semantical level which makes them more interpretable.&lt;/p&gt;

&lt;p&gt;Due to complex interactions within data which might yield biases or spurious correlations, doctors favour a more interpretable machine learning method as linear regression, although its performance is not as high as the deep learning methods. One example here:&lt;/p&gt;

&lt;p&gt;Caruana, R., Lou, Y., Gehrke, J., Koch, P., Sturm, M., &amp;amp; Elhadad, N. (2015). Intelligible Models for HealthCare. Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD ’15, 1721–1730. &lt;a href=&quot;https://doi.org/10.1145/2783258.2788613&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Some medical AI research suffer has the same issues present in the whole AI field: exaggerated claims (e.g. claiming to solve a task but solving solely a proxy sub-task), small datasets, evaluation issues. Particularly in medicine where the diseases are rare, taking into account prevalence helps assessing the robustness of a &lt;a href=&quot;https://spectrumnews.org/opinion/viewpoint/quest-autism-biomarkers-faces-steep-statistical-challenges/&quot;&gt;system&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;However, in 2017 we had a few breakthroughs in medical AI:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., … Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA - Journal of the American Medical Association, 316(22), 2402–2410. &lt;a href=&quot;https://doi.org/10.1001/jama.2016.17216&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Salehinejad, H., Valaee, S., Mnatzakanian, A., Dowdell, T., Barfett, J., &amp;amp; Colak, E. (2017). Interpretation of Mammogram and Chest X-Ray Reports Using Deep Neural Networks - Preliminary Results. &lt;a href=&quot;https://arxiv.org/abs/1708.09254&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Merkow, J., Lufkin, R., Nguyen, K., Soatto, S., Tu, Z., &amp;amp; Vedaldi, A. (2017). DeepRadiologyNet: Radiologist Level Pathology Detection in CT Head Images, 1–22. &lt;a href=&quot;https://arxiv.org/abs/1711.09313&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These papers propose large datasets and a solid evaluation which assumes a ROC sensitivity/specificity and a comparison with doctors. In contrast to other AI (e.g. self driving cars, recommendation systems), none of these medical AI has been implemented in hospitals as medical solutions.&lt;/p&gt;

&lt;h2 id=&quot;potential-for-confirmation-bias-in-personalized-explanations&quot;&gt;Potential for confirmation bias in personalized explanations&lt;/h2&gt;

&lt;p&gt;People express greater confidence in a hypothesis, although false, when asked to generate explanations for it.&lt;/p&gt;

&lt;p&gt;Koehler, D. J. (1991). Explanation, imagination, and confidence in judgment. Psychological Bulletin, 110(3), 499–519. &lt;a href=&quot;https://doi.org/10.1037//0033-2909.110.3.499&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Explanations:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Constrain inference by excluding possibilities inconsistent with current beliefs&lt;/li&gt;
  &lt;li&gt;Guide generalization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Explanations override:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Similarity&lt;/li&gt;
  &lt;li&gt;Diversity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Lombrozo, T. (2006). The structure and function of explanations. Trends in Cognitive Sciences, 10(10), 464–470. &lt;a href=&quot;https://doi.org/10.1016/j.tics.2006.08.004&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Explanations can be used to manipulate people’s decisions or to facilitate learning:&lt;/p&gt;

&lt;p&gt;Baumeister, R. F., &amp;amp; Newman, L. S. (1994). Self-Regulation of Cognitive Inference and Decision Processes. Personality and Social Psychology Bulletin, 20(1), 3–19. &lt;a href=&quot;https://doi.org/10.1177/0146167294201001&quot;&gt;pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Interpretability in machine learning advocates for human-centric explanations of black-box models. In simple tasks, the explanation is straightforward, for instance emphasizing the pixels corresponding to a detected object in an image. For other tasks which require slow judgements, a personalized explanation might enforce one’s confirmation biases, e.g. in topic detection a personalized interpretation underlines the words might point to one’s vocabulary. Therefore, in certain slow judgement tasks human-centric explanations as evaluation baselines for interpretability can be biased towards certain individuals.&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/interpretability/</guid>
			</item>
		
			<item>
				<title>Catastrophic forgetting</title>
				<link>https://nkundiushuti.github.io/blog/catastrophic-forgetting/</link>
				<pubDate>Sun, 28 Jan 2018 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;During October 2017 and December 2017 I have done an internship at Telefonica R&amp;amp;D department, working with &lt;a href=&quot;https://joanserra.weebly.com&quot;&gt;Joan Serra&lt;/a&gt;. It’s been 6 years since I collaborated with him, and I am happy to have worked again. One can learn a lot from Joan, particularly when it comes to his sound research methodology and writing skills. It was good to shift a bit from music and work on a more general topic in machine learning. This time it was catastrophic forgetting.&lt;/p&gt;

&lt;div class=&quot;alert-box text &quot;&gt;&lt;p&gt;J. Serra, D. Suris, M. Miron, and A. Karatzoglou, “Overcoming catastrophic forgetting with hard attention to the task” International Conference on Machine Learning 2018 (submitted)&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;A &lt;a href=&quot;https://arxiv.org/abs/1801.01423&quot;&gt;draft of the paper&lt;/a&gt; is already on arxiv and you can take a look.&lt;/p&gt;

&lt;p&gt;What is catastrophic forgetting?&lt;/p&gt;

&lt;p&gt;Ideally we want machines to learn in a sequential fashion, accumulating the knowledge learned in previous tasks, and using it to help future learning. In the more specific deep learning context, &lt;a href=&quot;https://www.cs.uic.edu/~liub/lifelong-machine-learning.html&quot;&gt;catastrophic forgetting&lt;/a&gt; is an undesired but natural behaviour of neural networks when trained to solve more than one task. By these means, catastrophic forgetting is related to other machine learning problems such as incremental learning and multitask learning.&lt;/p&gt;

&lt;p&gt;So what was the approach? Actually, it was Joan’s idea on introducing the id of the task as an input into the network, in other terms, putting hard attention to the task. In paralel to the network, the task id is encoded as a set of embeddings which after a gating mechanism become almost binary. This is multiplied element-wise with the activations of each layer in the network, protecting weights which were important for previous tasks, and, thus, not forgetting. Then, through a regularization parameter which enforces sparsity, the network can compress its capacity to the point of allowing the remaining capacity for the future tasks.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://nkundiushuti.github.io/images/hat.png&quot; alt=&quot;phenicx&quot; /&gt;&lt;/p&gt;

&lt;p&gt;I was glad to contribute on implementing or improving state of the art methods. It was real fun to have it a go with PathNet and Progressive Networks which were two approaches published by researchers in DeepMind.&lt;/p&gt;

&lt;p&gt;Then benchmark and the evaluation is much more exhaustive than the previous papers. It is done not on two or three, often very similar datasets/tasks, but on seven very different tasks.&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/catastrophic-forgetting/</guid>
			</item>
		
			<item>
				<title>Deep learning source separation for hip hop and classical music</title>
				<link>https://nkundiushuti.github.io/blog/ismir-and-mml/</link>
				<pubDate>Sun, 29 Oct 2017 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;During October I have attended the &lt;a href=&quot;https://musml.weebly.com&quot;&gt;Music and Machine Learning Workshop&lt;/a&gt; in Barcelona and the &lt;a href=&quot;https://ismir2017.smcnus.org&quot;&gt;ISMIR&lt;/a&gt; conference in Suzhou, China presenting papers on source separation for hiphop and Western classical music. You can check the PhD section on this website if you are new to source separation and you want to understand what I have been doing during the past 4 years.&lt;/p&gt;

&lt;p&gt;So what do hiphop and classical music have in common? :) Rather than having an universal model which separates any music piece, we were targeting context-specific approaches which deal with  problems which one encounters in particular music genres. For instance, in hiphop the vocal part is not sung and there is wide variety of timbres and production styles. On the other hand, in classical music, the complexity arrises from the multitude of harmonic instruments of similar timbre (depending on the piece). In this case, there are also opportunities which are given by the scores: the Western classical music pieces depart from symbolic representations.&lt;/p&gt;

&lt;p&gt;How can the deep learning or data-driven approaches take advantage of these characteristics? According to &lt;a href=&quot;https://www.deeplearningbook.org&quot;&gt;Ian Goodfellow&lt;/a&gt; a way to increase generalization is through data augmentation. So, hip hop and classical music source separation can improve using data augmentation or generation. This is the main idea behind the two papers.&lt;/p&gt;

&lt;p&gt;The MML paper, &lt;a href=&quot;https://mtg.upf.edu/node/3825&quot;&gt;Data augmentation for deep learning source separation of HipHop songs&lt;/a&gt; is based on work done by Hector Martel during his &lt;a href=&quot;https://repositori.upf.edu/bitstream/handle/10230/32919/Martel_2017.pdf?sequence=1&amp;amp;isAllowed=y&quot;&gt;undergrad thesis&lt;/a&gt; at UPF. He is also a hiphop producer so he proposed a &lt;a href=&quot;https://doi.org/10.5281/zenodo.823037&quot;&gt;hiphop dataset&lt;/a&gt;. You can check out the &lt;a href=&quot;https://hiphopss.github.io&quot;&gt;demo&lt;/a&gt; he did for his thesis’ presentation. I uploaded the presentation on &lt;a href=&quot;https://www.slideshare.net/MariusMiron2/presentation-mml&quot;&gt;slideshare&lt;/a&gt;. Here’s a video of the demo:&lt;/p&gt;

&lt;div class=&quot;flex-video&quot;&gt;
        &lt;iframe width=&quot;1280&quot; height=&quot;720&quot; src=&quot;https://www.youtube.com/watch?v=h8v0_3qKLHo&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&gt;&lt;/iframe&gt;
&lt;/div&gt;

&lt;p&gt;The ISMIR paper, &lt;a href=&quot;https://mtg.upf.edu/node/3806&quot;&gt;Monaural score-informed source separation for classical music using convolutional neural networks&lt;/a&gt; is a part of my PhD thesis on orchestral music source separation. It’s one of the first papers trying to improve deep learning source separation methods with score information. Similarly to the other deep learning papers, the code is made available through the &lt;a href=&quot;https://github.com/MTG/DeepConvSep&quot;&gt;github repository&lt;/a&gt;, and the separated tracks and computed evaluation metrics are on &lt;a href=&quot;https://doi.org/10.5281/zenodo.1009136&quot;&gt;zenodo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Demos:&lt;/p&gt;

&lt;div class=&quot;flex-video&quot;&gt;
        &lt;iframe width=&quot;1280&quot; height=&quot;720&quot; src=&quot;https://www.youtube.com/watch?v=c0xJIJrp5w8&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&gt;&lt;/iframe&gt;
&lt;/div&gt;

&lt;div class=&quot;flex-video&quot;&gt;
        &lt;iframe width=&quot;1280&quot; height=&quot;720&quot; src=&quot;https://www.youtube.com/watch?v=9vSxRVh1YZU&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&gt;&lt;/iframe&gt;
&lt;/div&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/ismir-and-mml/</guid>
			</item>
		
			<item>
				<title>Sound and Music Computing conference</title>
				<link>https://nkundiushuti.github.io/blog/smcconference/</link>
				<pubDate>Sun, 09 Jul 2017 00:00:00 +0000</pubDate>
				<description>&lt;p&gt;During July 5th and 9th I have attended the &lt;a href=&quot;https://smc2017.aalto.fi&quot;&gt;Sound and Music Computing Conference&lt;/a&gt; in Helsinki, Finland. I presented a &lt;a href=&quot;https://mtg.upf.edu/node/3765&quot;&gt;paper&lt;/a&gt; on &lt;a href=&quot;https://www.upf.edu/web/mtg/news/-/asset_publisher/WM181VyAQipW/content/id/16769628/maximized#.WV99oPy6yRs&quot;&gt;generating training data&lt;/a&gt; for deep learning source separation method, particularly in classical music, where you have the score but no multi-track data. The slides can be found &lt;a href=&quot;https://drive.google.com/file/d/0Bxgc1jYXBwD6cngwSWRLU2FEbVU/view?usp=sharing&quot;&gt;online&lt;/a&gt; and the code is on the source separation github repository &lt;a href=&quot;https://github.com/MTG/DeepConvSep&quot;&gt;DeepConvSep&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I had the opportunity to visit the acoustics lab at Aalto University and attend a few demos. I’ve been in the anechoic rooms where they recorded the orchestra dataset which I have annotated and used in my paper on &lt;a href=&quot;https://www.hindawi.com/journals/jece/2016/8363507/&quot;&gt;score-informed orchestral separation&lt;/a&gt;. Interestingly, in one of the demos, Jukka Patynen convolved close-microphone recordings with impulse responses taken from famous concert venues, to demonstrate how different the same recording can sound in varius halls.&lt;/p&gt;

&lt;p&gt;There were quite a few interesting posters, from which I mention:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Interesting app which doesn’t see the learning process as a game but rather as a collaborative process, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p77.pdf&quot;&gt;Burns et al. : Learning to Play the Guitar at the Age of Interactive and Collaborative Web&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Source separation for voice removal and then DTW for the automatic accompaniment, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p110.pdf&quot;&gt;Wada et al. : An Adaptive Karaoke System that Plays Accompaniment Parts of Music Audio Signals Synchronously with Users’ Singing Voices&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p104.pdf&quot;&gt;Kirkbride: Troop: A Collaborative Tool for Live Coding&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;From the presented papers, I was mostly interested in the MIR ones using deep learning, but there were also some other interesting ones:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Emotion recognition using many features and MFCCs with two CNN branches for each emotion, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p208.pdf&quot;&gt;Malik et al.: Stacked Convolutional and Recurrent Neural Networks for Music Emotion&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;In contrast to other architectures which use large filters, they use small filters, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p220.pdf&quot;&gt;Lee et al.: Sample-Level Deep Convolutional Neural Networks for Music Auto-Tagging Using Raw Waveforms&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Basically in-ear filtering and mixing for hearing protection, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p306.pdf&quot;&gt;Albrecht et al.: Electronic Hearing Protection for Musicians&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Meter detection in symbolic files, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p373.pdf&quot;&gt;Mcleod et al. : Meter Detection in Symbolic Music Using a Lexicalized PCFG&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Transcription with RNNs and alignment with DTW for piano recordings, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p380.pdf&quot;&gt;Kwon et al.: Audio-to-Score Alignment of Piano Music Using RNN-Based Automatic Music &lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Orchestration using a RBM with a binarized version of the piano roll as input, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p434.pdf&quot;&gt;Crestel et al.: Live Orchestral Piano, a System for Real-Time Orchestral Music Generation&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Active listening; they change onsets and pitches of the voice, &lt;a href=&quot;https://smc2017.aalto.fi/media/materials/proceedings/SMC17_p443.pdf&quot;&gt;Ojima et al.: A Singing Instrument for Real-Time Vocal-Part Arrangement of Music and Audio Signals&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The proceedings can be found &lt;a href=&quot;https://smc2017.aalto.fi/proceedings.html&quot;&gt;online&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I really liked the keynote from Anssi Klapuri. A few years ago he moved from academia to Yousician, a company that develops apps for music learning. He presented the app and the architecture of their audio engine (they use Unity for graphics but for audio they use their own audio engine). Interestingly, most of the audio part is minimally processed and the audio is rooted as soon as possible to the output. The cross-cancelling seemed quite robust. Also, I really liked the testing techniques: they record the impulse responses of different phones and instruments so they can use this data afterwards for testing.&lt;/p&gt;

</description>
				<guid isPermaLink="true">https://nkundiushuti.github.io/blog/smcconference/</guid>
			</item>
		
	</channel>
</rss>
