<node id="606842">
  <nid>606842</nid>
  <type>event</type>
  <uid>
    <user id="28475"><![CDATA[28475]]></user>
  </uid>
  <created>1528402732</created>
  <changed>1528402732</changed>
  <title><![CDATA[Ph.D. Dissertation Defense - Zhong Meng]]></title>
  <body><![CDATA[<p><strong>Title</strong><em>:&nbsp; </em><em>Discriminative and Adaptive Training for Robust Speech Recognition and Understanding</em></p>

<p><strong>Committee:</strong></p>

<p>Dr. Biing-Hwang Juang, ECE, Chair , Advisor</p>

<p>Dr. Chin-Hui Lee, ECE</p>

<p>Dr. Elliott Moore, ECE</p>

<p>Dr. James McClellan, ECE</p>

<p>Dr. Yao Xie, ISyE</p>

<p><strong>Abstract:</strong></p>

<p>Robust automatic speech recognition (ASR) and understanding (ASU) under noisy conditions remains to be a challenging problem even with the advances of deep learning.&nbsp;To achieve robust ASU, two discriminative training objectives are proposed for keyword spotting and topic classification: (1) To accurately recognize the semantically important keywords, the non-uniform error cost minimum classification error training of DNN and BLSTM acoustic models is proposed to minimize the recognition errors of only the keywords. (2)&nbsp;To compensate for the mismatched objectives of speech recognition and understanding, minimum semantic error cost training of the BLSTM acoustic model is proposed to generate semantically accurate lattices for topic classification.</p>

<p>Further, to expand the application of the ASU system to various conditions,&nbsp;four adaptive training approaches are proposed to&nbsp;improve the robustness of the ASR under different conditions: (1)&nbsp;To suppress the effect of inter-speaker variability on speaker-independent DNN acoustic&nbsp;model, speaker-invariant training is proposed to learn a deep representation in the DNN that is both senone-discriminative and speaker-invariant through adversarial multi-task training&nbsp;(2)&nbsp;To achieve condition-robust unsupervised adaptation with parallel data, adversarial teacher-student learning is proposed to suppress multiple factors of condition variability&nbsp;in the procedure of knowledge transfer from a well-trained source domain LSTM acoustic model to the target domain.&nbsp;(3)&nbsp;To further improve the adversarial learning for unsupervised adaptation with unparallel data, domain separation networks are used to enhance the domain-invariance of the&nbsp;senone-discriminative deep representation by explicitly modeling the private component that&nbsp;is unique to each domain. (4)&nbsp;To achieve robust far-field ASR, an LSTM adaptive beamforming network is proposed to estimate the real-time beamforming filter coefficients to cope with non-stationary environmental noise and dynamic nature of source and microphones positions.</p>
]]></body>
  <field_summary_sentence>
    <item>
      <value><![CDATA[Discriminative and Adaptive Training for Robust Speech Recognition and Understanding ]]></value>
    </item>
  </field_summary_sentence>
  <field_summary>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_summary>
  <field_time>
    <item>
      <value><![CDATA[2018-06-22T11:00:00-04:00]]></value>
      <value2><![CDATA[2018-06-22T13:00:00-04:00]]></value2>
      <rrule><![CDATA[]]></rrule>
      <timezone><![CDATA[America/New_York]]></timezone>
    </item>
  </field_time>
  <field_fee>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_fee>
  <field_extras>
      </field_extras>
  <field_audience>
          <item>
        <value><![CDATA[Public]]></value>
      </item>
      </field_audience>
  <field_media>
      </field_media>
  <field_contact>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_contact>
  <field_location>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_location>
  <field_sidebar>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_sidebar>
  <field_phone>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_phone>
  <field_url>
    <item>
      <url><![CDATA[]]></url>
      <title><![CDATA[]]></title>
            <attributes><![CDATA[]]></attributes>
    </item>
  </field_url>
  <field_email>
    <item>
      <email><![CDATA[]]></email>
    </item>
  </field_email>
  <field_boilerplate>
    <item>
      <nid><![CDATA[]]></nid>
    </item>
  </field_boilerplate>
  <links_related>
      </links_related>
  <files>
      </files>
  <og_groups>
          <item>434381</item>
      </og_groups>
  <og_groups_both>
          <item><![CDATA[ECE Ph.D. Dissertation Defenses]]></item>
      </og_groups_both>
  <field_categories>
          <item>
        <tid>1788</tid>
        <value><![CDATA[Other/Miscellaneous]]></value>
      </item>
      </field_categories>
  <field_keywords>
          <item>
        <tid>100811</tid>
        <value><![CDATA[Phd Defense]]></value>
      </item>
          <item>
        <tid>1808</tid>
        <value><![CDATA[graduate students]]></value>
      </item>
      </field_keywords>
  <field_userdata><![CDATA[]]></field_userdata>
</node>
