<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>PolMine Project</title>
    <description>Website of the PolMine Project</description>
    <link>http://polmine.github.io/</link>
    <atom:link href="http://polmine.github.io/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Fri, 23 May 2025 13:12:50 +0000</pubDate>
    <lastBuildDate>Fri, 23 May 2025 13:12:50 +0000</lastBuildDate>
    <generator>Jekyll v3.9.2</generator>
    
      <item>
        <title>GermaParl2: Constitution Day 2025 Release</title>
        <description>&lt;h2 id=&quot;germaparl2--constitution-day-2025-release&quot;&gt;GermaParl2 – Constitution Day 2025 Release&lt;/h2&gt;

&lt;p&gt;We are pleased to seize Germany’s 2025 Constitution Day as an opportunity for the release of the newest version of &lt;a href=&quot;https://zenodo.org/records/15495748&quot;&gt;GermaParl (v2.3.0-rc1)&lt;/a&gt;. The new version provides incremental quality improvements and covers all sessions of Germany’s Bundestag. The corpus includes 291 million tokens in 4559 protocols of the entire 20 legislative debates until March 18, 2025. The new GermaParl2 version is the up-to-date resource for researchers eager to analyse parliamentary debates during Germany’s “Ampel” government.&lt;/p&gt;

&lt;p&gt;We offer a beta today that is available for registered users. Prospective users can find more information about access to the data and data installation on &lt;a href=&quot;https://zenodo.org/records/15495748&quot;&gt;Zenodo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Developing and maintaining GermaParl2 is at the intersection of work within the context of the &lt;a href=&quot;https://www.nfdi.de/&quot;&gt;NFDI&lt;/a&gt; consortia &lt;a href=&quot;https://text-plus.org/&quot;&gt;Text+&lt;/a&gt; (updates, quality improvements, standardisation) and &lt;a href=&quot;https://www.konsortswd.de/&quot;&gt;KonsortSWD&lt;/a&gt; (linked data): Aside from presenting current updates, we also use this occasion to take stock of past developments and cast a glance into the future of GermaParl.&lt;/p&gt;

&lt;h2 id=&quot;towards-a-high-quality-full-coverage-resource&quot;&gt;Towards a High-Quality, Full Coverage Resource&lt;/h2&gt;

&lt;p&gt;The first version of GermaParl2 was released on May 23, Constitution Day, of 2023 and presented a substantive update of coverage compared to the initial version of GermaParl. &lt;a href=&quot;https://polmine.github.io/posts/2023/05/23/GermaParl2-Constitution-Day-Release.html&quot;&gt;GermaParl v2.0.0&lt;/a&gt; provides access to all debates between September 1949 and September 2021 in two formats which comprise rich structural and linguistical annotation and facilitate many kinds of analyses in social science research and beyond. The corpus was released on &lt;a href=&quot;https://zenodo.org/records/7949074&quot;&gt;Zenodo&lt;/a&gt; and accompanied by comprehensive &lt;a href=&quot;https://polmine.github.io/GermaParl2/&quot;&gt;documentation&lt;/a&gt; and continuous community outreach.&lt;/p&gt;

&lt;p&gt;Following-up this initial release, we provided four releases of GermaParl via Zenodo either to the broader public or as release candidates for which access can be granted upon request. &lt;a href=&quot;https://polmine.github.io/posts/2024/02/20/GermaParl2-Update.html&quot;&gt;GermaParl v2.0.1&lt;/a&gt;, released in December 2023, provided some incremental quality improvements of the initial release. Following about half a year in closed beta as GermaParl v2.1.0-rc2, GermaParl v2.1.0 was made available as a public release in July 2024. GermaParl v2.1.0 extended the temporal coverage of GermaParl2 to July 2023 and introduced date-specific assignments of party affiliations for the 20th legislative period. In earlier versions of the corpus, the limited availability of structured date-specific data on speakers’ party affiliations resulted in a rather coarse granularity of party affiliation assignments: We were not able to represent changes in party affiliations of speakers within a legislative period. So each speaker was assigned to the same party throughout a legislative period. However, this situation is changing, and more detailed information on party affiliation is becoming available in more structured formats. As a start, we added date-specific party affiliation information for Members of Parliament in the 20th legislative period (using detailed albeit mostly unstructured party affiliation data offered by the corresponding Wikipedia overview page, see &lt;a href=&quot;https://de.wikipedia.org/wiki/Liste_der_Mitglieder_des_Deutschen_Bundestages_(20._Wahlperiode)&quot;&gt;here&lt;/a&gt;), but emerging new resources potentially facilitate more nuanced annotations of party affiliations throughout the corpus in the future. This will be an important next step in the development of GermaParl.&lt;/p&gt;

&lt;h2 id=&quot;germaparl-as-linked-data&quot;&gt;GermaParl as Linked Data&lt;/h2&gt;

&lt;p&gt;With the release of GermaParl v2.2.0-rc1 as a release candidate in July 2024, the corpus was extended to cover all sessions until June 2024. More importantly from a technical point of view, a novel feature was introduced: The inclusion of Uniform Resource Identifiers (URIs) of the DBpedia Knowledge Graph for persons, organizations and locations in continuous text. Adding URIs to these entities in the Corpus Workbench (CWB) version of the corpus via the &lt;a href=&quot;https://www.dbpedia-spotlight.org&quot;&gt;DBpedia Spotlight&lt;/a&gt; Entity Linking tool is the first step toward GermaParl as a linked textual resource. Using the toolset developed by us in the project “Linking Textual Data” (as part of KonsortSWD within the National Research Data Infrastructure/NFDI), in particular the R package &lt;a href=&quot;https://github.com/PolMine/dbpedia&quot;&gt;dbpedia&lt;/a&gt;, we were able to add these URIs as a structural attribute to the CWB version of GermaParl.&lt;/p&gt;

&lt;p&gt;The addition of URIs is also part of the newest release of the corpus, GermaParl v2.3.0-rc1 which we announce today as a closed-beta release. While the inclusion of entity-specific URIs is currently still experimental and its quality not yet checked, the potentials of URIs for substantive research are manifold – for example allowing the disambiguation or enrichment of entities. By making this new annotation layer available for the GermaParl user community at an early stage, we want to foster discussion and advance the development of tools and data in a community-driven fashion.&lt;/p&gt;

&lt;p&gt;Using the new annotation layer should be easy: Aside from the new release candidate of GermaParl, anything you need to get started is the following CQP query syntax which will allow you to look for regions described by a specific URI in the new structural attribute “dbpedia_uri”:&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;s1&quot;&gt;&apos;/region[dbpedia_uri,a]::a.dbpedia_uri=&quot;http://de.dbpedia.org/resource/Europa&quot;&apos;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;This query syntax can be used in places which allow CQP queries – such as the methods of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;count()&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;kwic()&lt;/code&gt; in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;polmineR&lt;/code&gt; R package. This can be useful in instances in which the same concept is represented by a number of different expressions or when different concepts are described by the same words.&lt;/p&gt;

&lt;p&gt;For a first impression, we could look at sequences of words which the query above corresponds to:&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;GERMAPARL2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;/region[dbpedia_uri,a]::a.dbpedia_uri=&quot;http://de.dbpedia.org/resource/Europa&quot;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cqp&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;TRUE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;breakdown&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;TRUE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot; data-lang=&quot;text&quot;&gt;## Error in count(&quot;GERMAPARL2&quot;, query = &quot;/region[dbpedia_uri,a]::a.dbpedia_uri=\&quot;http://de.dbpedia.org/resource/Europa\&quot;&quot;, : konnte Funktion &quot;count&quot; nicht finden&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;For a few more details on the annotation process, associated potentials and challenges as well as future steps, we refer to the “Cookin’ with GermaParl” webinar series.&lt;/p&gt;

&lt;h2 id=&quot;germaparl-v230-rc1-high-quality-coverage-entire-20th-legislative-period&quot;&gt;GermaParl v2.3.0-rc1: High-Quality Coverage, Entire 20th Legislative Period&lt;/h2&gt;

&lt;p&gt;After GermaParl v2.2.0-rc1 extended the temporal coverage of the corpus to June 2024, the most recent update, GermaParl v2.3.0-rc1 which we release today adds all remaining protocols of the 20th legislative period until March 2025. It makes use of the date-specific assignments of party affiliations introduced with the release of GermaParl v2.1.0. In more recent debates, this modification facilitates the date-specific differentiation between members of the “DIE LINKE” and members of the “Bündnis Sahra Wagenknecht” (“BSW”) in particular. Please note that the assignment of date-specific party affiliations is currently still limited to the 20th legislative period. Related to this, the difference between “DIE LINKE” and “Die Linke” as a parliamentary group is deliberate and indicates the difference between parliamentary group (“Fraktion”) and a group with group status after the parliamentary group split up (“Gruppe”).&lt;/p&gt;

&lt;p&gt;Users familiar with GermaParl v2.2.0-rc1 might notice some additional changes. In an attempt to further improve the quality of the data, we consolidated the full names and party affiliations we assigned to speakers extracted from the protocols. While consistency was always an important motivation, in some instances, the same speaker could be assigned to different variations of the same name (e.g., including or omitting middle initials) in different legislative periods. In other instances, a speaker could be assigned to a party in one parliamentary role but not in another due to the different resources used to enrich different speaker roles. GermaParl v2.3.0-rc1 makes an effort to remedy the most obvious instances of contradictory metadata assignments. Other changes include the recoding of the abbreviation for the “Zentrumspartei” from “Z” (which is used in the original protocols to indicate the respective parliamentary group) to “DZP” (for “Deutsche Zentrumspartei”) which is the preferred abbreviation of the party in the “Stammdaten” file, a collection of metadata of Members of Parliaments provided by the German Bundestag. We also consolidated the party affiliation of “Ludwig Erhard”. After indicating in earlier versions that the party affiliation ultimately seems to remain unclear, we follow the general assumption expressed in both the Wikipedia overview pages and in other resources such as the “Parliaments Day-by-Day” database by Turner-Zwinkels and colleagues (2022) and assign “CDU” throughout.&lt;/p&gt;

&lt;h2 id=&quot;where-we-stand-looking-ahead&quot;&gt;Where We Stand, Looking Ahead&lt;/h2&gt;

&lt;p&gt;After the release of GermaParl v2.2.0-rc1 constituted an extensive update of the resource, GermaParl v2.3.0-rc1 provides further incremental improvements and extends the coverage of the corpus to March 2025. Now, GermaParl comprises the first 20 legislative periods in their entirety. As before, the development GermaParl relies on community involvement: With the experimental nature of new features, user feedback is invaluable for the future development of the resource. While we are convinced that the addition of Uniform Resource Identifiers greatly enhances the usefulness of GermaParl, specific conceptual, methodological and technical decisions should be made with usability and accessibility in mind. In addition, the quality of these new annotations is not systemically evaluated. With this release as a release candidate, we want to enable the community to contribute to the development of the resource by engaging in the discussion about the specific implementation and potential use cases. Extensive feedback constitutes an important part of our strategy to develop useful tools and data. Please reach out to us via email (dennis.schuele@uni-due.de) to report issues and suggestions.&lt;/p&gt;

&lt;p&gt;GermaParl v2.3.0-rc1 is also a transitionary release. For the preparation of this version, we reevaluated the way we represent speaker metadata. In the current release, we did not yet implement this to the full extent. Among others, changes of speaker names between or within legislative periods, for example due to marriage, are not yet properly represented. A better solution – fine-grained, date-specific annotations of names and party affiliations for more speakers – is on the horizon. To this end, a next major milestone is the provision of GermaParl as a resource in the XML encoding standard of the &lt;a href=&quot;https://www.clarin.eu/parlamint&quot;&gt;ParlaMint&lt;/a&gt; project. This new development will not only make it easier to represent additional, fine-grained and extensively documented metadata on speaker level but, crucially, will strengthen the interoperability of the resource, making comparative research more accessible and the integration of workflows and tools more seamless. By making use of more consolidated representations of metadata, GermaParl v2.3.0-rc1 is a first step toward this goal.&lt;/p&gt;

&lt;h2 id=&quot;acknowledgements&quot;&gt;Acknowledgements&lt;/h2&gt;

&lt;p&gt;We gratefully acknowledge funding by &lt;a href=&quot;https://www.konsortswd.de/&quot;&gt;KonsortSWD&lt;/a&gt; and &lt;a href=&quot;https://text-plus.org/&quot;&gt;Text+&lt;/a&gt; within the German National Research Data Infrastructure (&lt;a href=&quot;https://www.nfdi.de/&quot;&gt;Nationale Forschungsdateninfrastruktur/NFDI&lt;/a&gt;) to prepare, enrich and maintain the corpus and associated resources. We also want to thank the Institute of Contemporary History in Ljubljana, Slovenia, for the opportunity to work toward the ParlaMint version of GermaParl during a Visiting Fellowship in October 2024.&lt;/p&gt;

&lt;h2 id=&quot;technical-note&quot;&gt;Technical Note&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The release of GermaParl v2.3.0-rc1 comprises of three objects which are available on Zenodo. GermaParl v2.3.0-rc1 (CWB), GermaParl v2.3.0-rc1 (XML) and GermaParl v2.3.0-rc1.1 (XML). GermaParl v2.3.0-rc1 (XML) and GermaParl v2.3.0-rc1.1 (XML) are functionally identical except for slight improvements and the harmonization of agenda item annotations in legislative periods 19 and 20 in the latter. Since we do not include agenda items in the CWB version of the corpus due to the limited robustness of this annotation, both sets of XML files would result in the same CWB corpus. In consequence, we do not provide a separate CWB version for this set of XML files. Users interested in the CWB version (for example to be used with the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;polmineR&lt;/code&gt; R package) can use the GermaParl v2.3.0-rc1 CWB version which contains the described improvements and coverage. For users who are interested in the XML version, it is advised to use GermaParl v2.3.0-rc1.1. Although substantively mostly identical, there is little reason to use GermaParl v2.3.0-rc1 in its XML representation. We include GermaParl v2.3.0-rc1 (XML) for reasons of research data management and transparent versioning as it is this precise set of XML files which was used for the preparation of the CWB corpus.&lt;/em&gt;&lt;/p&gt;
</description>
        <pubDate>Fri, 23 May 2025 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2025/05/23/GermaParl2-Update.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2025/05/23/GermaParl2-Update.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>Update for GermaParl2 – Improving Corpus Quality and a Look Ahead</title>
        <description>&lt;h1 id=&quot;update-for-germaparl2--improving-corpus-quality-and-a-look-ahead&quot;&gt;Update for GermaParl2 – Improving Corpus Quality and a Look Ahead&lt;/h1&gt;

&lt;p&gt;We always envisioned GermaParl as an evolving resource. Since some issues only become apparent during productive work, the continuous provision of releases aimed at improving the corpus was always part of our roadmap. Accordingly, over the last months, we put the corpus to work in substantive analyses, comprehensive quality checks and educational outreach. This enabled us to spot remaining flaws. The same is true for users of the resource who might encounter bugs and missing features – some of which were brought to our attention via our &lt;a href=&quot;https://github.com/PolMine/GermaParl2/issues&quot;&gt;issue tracker&lt;/a&gt; on GitHub.&lt;/p&gt;

&lt;p&gt;As a first step towards even better data quality, we can gladly announce the first patch release for GermaParl2 today. Like the previous version, we made the corpus available via &lt;a href=&quot;https://zenodo.org/records/10416536&quot;&gt;Zenodo&lt;/a&gt;. The release of GermaParl v2.0.1 builds on the release version of GermaParl2, including the same features and covering the same period of time, but addresses a lot of the recently identified issues and provides additional improvements regarding the general quality of the resource. We want to highlight updates in three areas which are especially noteworthy:&lt;/p&gt;

&lt;h2 id=&quot;removed-appendices&quot;&gt;Removed appendices&lt;/h2&gt;

&lt;p&gt;We noticed that for quite a large number of sessions, the processed protocols not only contained speeches delivered on the plenary floor but also appendices (see &lt;a href=&quot;https://github.com/PolMine/GermaParl2/issues/1&quot;&gt;issue #1&lt;/a&gt; on GitHub). These comprise of different elements, in particular speeches which were only added to the minutes of the session. Their accidental inclusion was caused by a greater-than-expected variation in end-of-speech expressions. These appendices are now removed from the corpus.&lt;/p&gt;

&lt;h2 id=&quot;improved-speaker-recognition&quot;&gt;Improved speaker recognition&lt;/h2&gt;

&lt;p&gt;While most speakers are correctly identified, despite our best efforts, speakers are not properly recognized in every case. This was also noticed by members of the GermaParl community (see &lt;a href=&quot;https://github.com/PolMine/GermaParl2/issues/2&quot;&gt;issue #2&lt;/a&gt; on GitHub). At least for presidential speakers, we were able to improve the identification of speakers by introducing additional line breaks before presidential speakers start to speak. This enables our regex-based approach to recognize speaker calls properly even if line breaks are missing in the original data. In addition, the identification of speakers of the federal council was improved by adding and tuning regular expressions for this group of speakers.&lt;/p&gt;

&lt;h2 id=&quot;large-paragraphs-in-lps-13-to-18&quot;&gt;Large Paragraphs in LPs 13 to 18&lt;/h2&gt;

&lt;p&gt;From earlier iterations of the corpus preparation, we already knew that the reconstruction of paragraphs and the concatenation of stage expressions – i.e., elements interrupting a speaker’s utterance – can be limited by noise which is mostly introduced by issues in the raw data. In some cases, this results in unexpected behavior such as the unintended concatenation of multiple lines which can obscure valid speaker calls. While this has been addressed for earlier legislative periods in the initial release of GermaParl, the changing nature of the raw data unfortunately lead to remaining overly large paragraphs in legislative periods 13 to 18. This is improved now, resulting in additionally identified speakers. Especially protocols in legislative periods 15 and 16 benefit from this improvement. 
Furthermore, several minor but meaningful improvements are included in this release.  See the change log in the &lt;a href=&quot;https://polmine.github.io/GermaParl2/&quot;&gt;documentation&lt;/a&gt; for all changes.&lt;/p&gt;

&lt;h2 id=&quot;germaparl-v210-release-candidate-2&quot;&gt;GermaParl v2.1.0 Release Candidate 2&lt;/h2&gt;

&lt;p&gt;Aside from this patch release, we want to use this opportunity to announce the closed-beta release of the next version of GermaParl2 - GermaParl v2.1.0-rc2. This release candidate includes all improvements of GermaParl v2.0.1 described above plus the first 116 protocols of the 20th legislative period. So, while GermaParl v2.0.1 (like the initial GermaParl2 release) covered parliamentary debates between September 1949 and September 2021, this new beta release extends this period to the time between September 1949 and July 2023. Unlike the public release of GermaParl v2.0.1, this update is provided via Zenodo with restricted access. While we are confident that the corpus preparation workflow allows us to create qualitative and reliable versions of the corpus, final quality checks which were performed for the previous corpus versions are still pending for the recently added protocols of the 20th legislative period. Releasing the upcoming version of GermaParl2 in this restricted way allows us to share the more recent debates in a level of quality which should be sufficient for a lot of purposes but is not yet fully checked for potentially remaining flaws. Accordingly, interested users should be aware of the potentially preliminary status of the resource. In particular, the restricted release makes it possible to include the community in this corpus curation effort: We invite interested users to apply for access on Zenodo and help us to further improve the resource before the final open release by reporting bugs via our &lt;a href=&quot;https://github.com/PolMine/GermaParl2/issues&quot;&gt;issue tracker&lt;/a&gt;. Requesting access is described on the corresponding &lt;a href=&quot;https://zenodo.org/records/10421773&quot;&gt;Zenodo page&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;next-steps&quot;&gt;Next Steps&lt;/h2&gt;

&lt;p&gt;GermaParl v2.0.1 addresses many issues which will be indicated as „closed“ in the issue tracker on GitHub. However, you will find that this update does not address all reported issues. In addition, new issues will certainly become apparent as people continue to use the resource. Please note that these issues might indicate current limitations of the resource for your particular use case. If the possibility to patch the resource becomes apparent, we will try to include it in an upcoming release. This particularly applies to the beta-release of GermaParl v2.1.0-rc2 which presents a good opportunity to address additional issues and potential feature requests.&lt;/p&gt;

&lt;p&gt;Finally, we want to re-iterate our suggestion to share your experiences - bugs, flaws, uncertainties – with us via GitHub issues. With GermaParl v2.1.0 already on the horizon, there will be upcoming releases and we greatly benefit from your feedback to further improve the resource.&lt;/p&gt;
</description>
        <pubDate>Tue, 20 Feb 2024 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2024/02/20/GermaParl2-Update.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2024/02/20/GermaParl2-Update.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>GermaParl2 Constitution Day Release 2023</title>
        <description>&lt;h1 id=&quot;germaparl2-constitution-day-release-2023&quot;&gt;GermaParl2 Constitution Day Release 2023&lt;/h1&gt;

&lt;p&gt;We are pleased to announce the public release of GermaParl v2.0.0 corpus today, on Germany’s Constitution Day (May 23, 2023). With GermaParl2, all parliamentary debates of the German Bundestag from 1949 to 2021 become available in a comprehensively annotated format.  The public release follows a process of two beta releases in which registered beta users provided valuable feedback on data quality and usability. In taking up these suggestions, we are confident that we are now able to offer a high-quality corpus. This is where the corpus is now available without the need to register. It is available under a Creative Commons license (&lt;a href=&quot;https://creativecommons.org/licenses/by-sa/4.0/&quot;&gt;CC BY-SA 4.0&lt;/a&gt;):&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;The data can be downloaded persistently from &lt;a href=&quot;https://zenodo.org/record/7949074&quot;&gt;Zenodo&lt;/a&gt;. Available data formats are XML, and an linguistically annotated and indexed version (Corpus Workbench / CWB): &lt;a href=&quot;https://zenodo.org/record/7949074&quot;&gt;https://zenodo.org/record/7949074&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;The XML version of GermaParl2 is also available at GitHub: &lt;a href=&quot;https://github.com/PolMine/GermaParlTEI&quot;&gt;https://github.com/PolMine/GermaParlTEI&lt;/a&gt;&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;GermaParl2 has been prepared as high-quality data for trustworthy scientific publications. As a matter of transparency, we offer extensive documentation. It includes a presentation of the data, its preparation workflow, and some more technical details as well as acknowledgements and remarks to future work. The documentation of the corpus is available &lt;a href=&quot;https://polmine.github.io/GermaParl2/&quot;&gt;here&lt;/a&gt;: &lt;a href=&quot;https://polmine.github.io/GermaParl2/&quot;&gt;https://polmine.github.io/GermaParl2/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;GermaParl2 covers the entire post-war history of German parliamentarism up the end of the 19th legislative period. It includes a total of 273550607 tokens. To offer a first glimpse into the data, we may learn from this barplot that the long-term trend is that more and more words are spoken in parliament per legislative period:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/2023-05-23-GermaParl2-Constitution-Day-Release/unnamed-chunk-2-1.png&quot; alt=&quot;plot of chunk unnamed-chunk-2&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Aside from coverage, comprehensive annotation is a unique feature of GermaParl2. The structure of the data facilitates the segmentation of debates into single utterances enriched with additional metadata. See the following table on the structural annotation of GermaParl2 and how it differs from GermaParl v1.0.6.&lt;/p&gt;

&lt;table&gt;
&lt;caption&gt;Mapping table for structural attributes (GermaParl/GermaParl2)&lt;/caption&gt;
 &lt;thead&gt;
  &lt;tr&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; GermaParl v1.0.6 &lt;/th&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; GermaParl v2.0.0-beta3 &lt;/th&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; Description &lt;/th&gt;
  &lt;/tr&gt;
 &lt;/thead&gt;
&lt;tbody&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; the main document node attributes &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; lp &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_lp &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; legislative period &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; session &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_no &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; session number &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; date &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_date &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; date in YYYY-MM-DD &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; year &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_year &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; year in YYYY &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; url &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_url &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; the url of the source document &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; src &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_filetype &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; the type of the source document &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; the main speaker node attributes &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_who &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; mostly the raw speaker call found in the protocols &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_name &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; consolidated speaker names &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; parliamentary_group &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_parlgroup &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; consolidated parliamentary group affiliation of a speaker &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; party &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_party &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; consolidated party affiliation of a speaker &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; role &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_role &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; role of a speaker (e.g. member of parliament, governmental actor, etc.) &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; p &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; a paragraph node, technically parent of sentence nodes &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; interjection &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; p_type &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; the type of paragraph, used to indicate paragraphs which are not speech (such as interjections) &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; s &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; a sentence node &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; ne &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; a named entity node to represent (nested) named entities &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; - &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; ne_type &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; the type of named entity &lt;/td&gt;
  &lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;As described in the &lt;a href=&quot;https://polmine.github.io/posts/2023/04/03/GermaParl-v2-beta3-Release-Note.html&quot;&gt;release note of the previous beta release&lt;/a&gt;, the labels of the attributes purposefully deviate from those of GermaParl v1 to better represent the hierarchical structure of the corpus. See also section “data formats” in this release note.&lt;/p&gt;

&lt;p&gt;The CWB version of GermaParl2 includes linguistic annotation layers. Aside from tokenization and sentence segmentation as well as named entities which are encoded as structural attributes, the linguistic annotation comprises of Part-of-Speech annotation (providing POS-Tags in both the Stuttgart-Tübingen-Tagset and the UD-Tagset) and lemmatization. For tokenization, sentence segmentation, Part-of-Speech annotation with the UD tagset and Named Entity Recognition, &lt;a href=&quot;https://stanfordnlp.github.io/CoreNLP/index.html&quot;&gt;Stanford CoreNLP&lt;/a&gt; is used. POS-Tags in the Stuttgart-Tübingen-Tagsets and Lemmata are added using the &lt;a href=&quot;https://www.cis.uni-muenchen.de/~schmid/tools/TreeTagger/&quot;&gt;TreeTagger&lt;/a&gt;.&lt;/p&gt;

&lt;h3 id=&quot;available-data-formats&quot;&gt;Available Data Formats&lt;/h3&gt;

&lt;p&gt;GermaParl2 is provided in two formats:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;An XML format which is inspired by the TEI-Standard&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;A CWB corpus based on these TEI-XML files&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The XML version of the corpus which is available on &lt;a href=&quot;https://github.com/PolMine/GermaParlTEI&quot;&gt;GitHub&lt;/a&gt; is designed to provide the data as a persistent, interoperable format. XML is not necessarily the first choice for fast analyses of corpus data. We provide it as an interoperable data format for users that may have their own pipelines for processing large-scale linguistic data. While the structural annotation of the data – i.e. the segmentation of the text into individual utterances, etc. – is provided in the XML version of the data, this version is not linguistically annotated.&lt;/p&gt;

&lt;p&gt;The &lt;a href=&quot;https://cwb.sourceforge.io/&quot;&gt;Corpus Workbench (CWB)&lt;/a&gt;]is an efficient corpus management tool for large text corpora. For the CWB version of the corpus, the XML files are imported into the Corpus Workbench. During this process, some additional steps to harmonize metadata on speaker level are performed to further increase the usability of the CWB corpus. A tarball with the CWB version of GermaParl2 is available from &lt;a href=&quot;https://zenodo.org/record/7949074&quot;&gt;Zenodo&lt;/a&gt; and can be used either by the CWB command line interface, graphical user-interfaces such as &lt;a href=&quot;https://cwb.sourceforge.io/cqpweb.php&quot;&gt;CQPweb&lt;/a&gt; or from within R, using the &lt;a href=&quot;https://CRAN.R-project.org/package=polmineR&quot;&gt;polmineR&lt;/a&gt; R package. The remainder of this release note focusses on the usage of GermaParl2 with R using polmineR.&lt;/p&gt;

&lt;h3 id=&quot;getting-started&quot;&gt;Getting Started&lt;/h3&gt;

&lt;p&gt;To use GermaParl2 with polmineR, the first step is to install the corpus. This should be as easy as…&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;install.packages&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;err&quot;&gt;“&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cwbtools&lt;/span&gt;&lt;span class=&quot;err&quot;&gt;”&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; 
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cwbtools&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;::&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus_install&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;doi&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;10.5281/zenodo.7949074&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;This will download the corpus and store it in an appropriate directory on your system. Depending on whether a CWB corpus has been installed before on the system, a dialogue will guide you through the process. If the previous beta version of GermaParl2 is already installed, the installation routine will suggest replacing the beta version with the new release. The downloaded data file is quite large, so depending on the available internet connection, this might take a few minutes.&lt;/p&gt;

&lt;p&gt;After installing all necessary packages and the corpus itself, the corpus should be available:&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;polmineR&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; 
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# should include GERMAPARL2 &lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;size&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;err&quot;&gt;“&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;GERMAPARL2&lt;/span&gt;&lt;span class=&quot;err&quot;&gt;”&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# test the functionality &lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;To increase the accessibility of the resource, we would like to offer the possibility of a rather informal exchange in the form of virtual cooking sessions (via Zoom). At the beginning of each lesson, we – members of the PolMine team – present a small recipe or use case involving polmineR and GermaParl2. These use cases should provide users with useful code snippets and should serve as an ice breaker for the subsequent Q&amp;amp;A and discussion.&lt;/p&gt;

&lt;p&gt;The first &lt;em&gt;Cookin’ with GermaParl&lt;/em&gt; session will address typical first analytical scenarios with polmineR and GermaParl2 and will take place on&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;June 1, 2023, 13:00 – 14:00&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you want to participate in these Cookin’ with GermaParl sessions, please send an email to stine.ziegler@uni-due.de.&lt;/p&gt;

&lt;h3 id=&quot;what-is-new&quot;&gt;What is new?&lt;/h3&gt;

&lt;p&gt;GermaPar2 has benefitted substantially from the feedback we received from beta users. The final release version of GermaParl2 differs in three substantial aspects from the previous beta releases:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;Increased data quality&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Genuine preparation of the debates included in GermaParl v1&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Changed enrichment of the speaker metadata&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id=&quot;increased-data-quality&quot;&gt;Increased Data Quality&lt;/h4&gt;

&lt;p&gt;After some additional steps of quality control and user feedback, further improvements in data quality were implemented. In particular, the detection of procedural comments. In the XML version, these are annotated as &lt;stage&gt; nodes, in the CWB corpus, these are annotated as paragraphs of type “stage” – were improved by reducing the number of false positives. This not only results in fewer text segments being falsely identified as comments, but also allows for the identification of more separate speeches which were previously hidden. Moreover, the refinement of regular expressions allowed for the identification of additional speakers. Several minor improvements include an improved way to detect and remove boilerplate text from the protocols (headers, for example) and more complete metadata on document level.&lt;/stage&gt;&lt;/p&gt;

&lt;h4 id=&quot;fresh-preparation-of-legislative-periods-1318&quot;&gt;Fresh preparation of legislative periods 13–18&lt;/h4&gt;

&lt;p&gt;The previous version of GermaParl (GermaParl v1) covered legislative periods 13 to 18. Until now, this data was incorporated into GermaParl2 by using the TEIs which have been prepared for GermaParl v1, making some minor adjustments such as fixing metadata and speaker attributes where necessary. This was done to ensure that the high data quality of GermaParl v1 is included in GermaParl2.&lt;/p&gt;

&lt;p&gt;However, this approach has one drawback. Using the structurally annotated TEI documents limits the possibility for comprehensive changes to the structure. In particular, the identification of additional speakers which might have been missed in the previous version is difficult when reusing the previous TEIs. Furthermore, it adds a layer of complexity to the data preparation process, limiting reproducibility and clean documentation.&lt;/p&gt;

&lt;p&gt;For the final version of GermaParl2, a new approach was used that should retain the quality of the existing data while allowing further improvements of the data. Instead of reusing the existing TEIs, the raw text was processed again from scratch making use of existing workflows instead of existing data. The result is very similar to GermaParl v1. However, additional speakers were identified – in particular speakers of the federal council but also some speakers such as the “Alterspräsident”. The inclusion of the 50th session of the 15th legislative period and the 135th session of the 18th legislative period which were missing previously are some major improvements aside from some minor changes to fill gaps in the document metadata, etc.&lt;/p&gt;

&lt;p&gt;It can be noted that the new version of the period covered by GermaParl v1 does differ in some additional aspects such as the reconstruction of paragraphs. Slight differences are expected.&lt;/p&gt;

&lt;h3 id=&quot;enrichment-of-the-speaker-metadata&quot;&gt;Enrichment of the speaker metadata&lt;/h3&gt;

&lt;p&gt;On speaker level, GermaParl2 like its predecessor, includes not only data which can be found in the protocols themselves – such as a speakers (family) name or the corresponding affiliation to a parliamentary group – but also additional information such as the party affiliation and a consolidated name of the speaker. In previous versions, the source of this additional information was mostly Wikipedia. For Members of Parliament this has changed. For the most part, data from the Stammdaten-File of the German Bundestag (available here)[https://www.bundestag.de/services/opendata] is used. It provides a host of additional information. However, since a speaker’s party affiliation are static in this Stammdaten file – each speaker has exactly one party assigned –, for GermaParl2 the Stammdaten are enriched with party assignments retrieved from Wikipedia. We mainly add the full name of the speaker and these party affiliations to the textual data. The switch should make the assignment of this additional information more transparent and reproducible and should result in more consistent annotations. It is planned to make the dataset available as an R package soon.&lt;/p&gt;

&lt;p&gt;Regarding this enrichment of speaker data, there are some other improvements in how the parliamentary data and the external data is matched. In addition, some minor harmonization of attributes such as the names of parties and parliamentary groups was applied to increase the usability of the resource.&lt;/p&gt;

&lt;h3 id=&quot;next-steps&quot;&gt;Next Steps&lt;/h3&gt;

&lt;p&gt;The public release of GermaParl2 is an important milestone in the evolution of the PolMine project. But there is more to come. As mentioned in the previous release note, a new XML version of the corpus is already under preparation. It moves the TEI-like XML we use now to a truly interoperable format based on the ParlaMint standard.&lt;/p&gt;

&lt;p&gt;If you wish to stay informed about the latest developments regarding GermaParl, you can subscribe to our project newsletter by emailing stine.ziegler@uni-due.de.&lt;/p&gt;

&lt;h3 id=&quot;feedback-welcome&quot;&gt;Feedback welcome!&lt;/h3&gt;

&lt;p&gt;We are confident that the data quality of the current release version is high. However, we are certain that further improvements in data quality are still achievable as soon as more users put the data to active use. In this sense, GermaParl v2.1.0 is already on the horizon. With GermaParl v2.0.0 as a solid foundation, improving data quality further should become a community effort. The best way to contribute to this endeavor is by reporting bugs and flaws in the data as well as suggestions and feature requests for further releases. One way to get in touch in this regard is by creating so called issues on (GitHub)[https://docs.github.com/en/issues/tracking-your-work-with-issues/creating-an-issue] in the (GermaParl2 repository)[https://github.com/PolMine/GermaParl2]. As mentioned above, this repository contains the TEI version of the current GermaParl corpus and thus should work as the hub for collecting feedback for both the XML and the CWB version of the corpus as the latter is based on the former.&lt;/p&gt;

&lt;p&gt;While GitHub issues allow users and developers of the resource to collaborate, discuss and evaluate different solutions interactively, reporting bugs and suggestions is of course also possible by email. If you prefer contributing by mail, please contact stine.ziegler@uni-due.de with any suggestions you might have.&lt;/p&gt;

&lt;p&gt;Aside from suggesting improvements for the data itself, we released the preparation workflow of the data. In the corresponding repository, suggestions are welcome as well.&lt;/p&gt;

&lt;h3 id=&quot;acknowledgments&quot;&gt;Acknowledgments&lt;/h3&gt;

&lt;p&gt;The data quality of GermaParl we are able to offer at this stage has benefitted significantly from a cooperation with the &lt;a href=&quot;https://www.uni-hildesheim.de/soldisk/&quot;&gt;SOLDISK&lt;/a&gt; project at the University of Hildesheim, and comprehensive manual quality control of the data carried out by the SOLDISK team. A very special thanks goes to Hannes Schammann, Max Kisselew, Franziska Ziegler, Carina Böker, Jennifer Elsner and Carolin McCrea.&lt;/p&gt;

&lt;p&gt;We also would like to thank all beta users for their invaluable feedback.&lt;/p&gt;

&lt;h3 id=&quot;funding-statement&quot;&gt;Funding statement&lt;/h3&gt;

&lt;p&gt;We gratefully acknowledge funding from the &lt;a href=&quot;https://www.nfdi.de/?lang=en&quot;&gt;German National Research Data Infrastructure (Nationale Forschungsdateninfrastruktur / NFDI)&lt;/a&gt;. Funding from &lt;a href=&quot;https://www.konsortswd.de&quot;&gt;KonsortSWD&lt;/a&gt; has advanced the data preparation tool set to facilitate the robust annotation of additional annotation layers in large corpora (such as Named Entities). This is instrumental for linking parliamentary data with other data. &lt;a href=&quot;https://www.konsortswd.de&quot;&gt;KonsortSWD&lt;/a&gt; is funded by the German Research Foundation (DFG) as part of the National Research Data Infrastructure Germany (Nationale Forschungsdateninfrastruktur, NFDI) under project number 442494171.&lt;/p&gt;

&lt;p&gt;Funding from the &lt;a href=&quot;https://www.text-plus.org&quot;&gt;Text+&lt;/a&gt; consortium is instrumental for updates of the corpus, quality control and keeping data formats up with current and future developments. &lt;a href=&quot;https://www.text-plus.org&quot;&gt;Text+&lt;/a&gt; is funded by the German Research Foundation (DFG) as part of the NFDI under project number 460033370.&lt;/p&gt;
</description>
        <pubDate>Tue, 23 May 2023 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2023/05/23/GermaParl2-Constitution-Day-Release.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2023/05/23/GermaParl2-Constitution-Day-Release.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>GermaParl v2.0.0-beta.3 Release Note</title>
        <description>&lt;h2 id=&quot;a-new-germaparl-v2-beta-version-to-improve-usability&quot;&gt;A new GermaParl v2 beta version to improve usability&lt;/h2&gt;

&lt;p&gt;On May 23 last year (Germany’s Constitution Day), we released the first beta version of GermaParl v2, a major rework of the GermaParl Corpus of Plenary Protocols of the German Bundestag. The most obvious development is that v2 comprises all protocols of plenary sessions in the German Bundestag (1949 - 2021). But there is also a change of the data structure. The “flat” scheme of structural attributes of GermaParl v1 had its merits for the usability of the data, but increasingly brought limitations. It inhibits performance and flexibility to handle nested attributes and annotations. The growth of the corpus as well as the inclusion of further annotation layers (sentence annotation and named entities to start with) induced us to move to a more hierarchical representation of the data with GermaParl v2.&lt;/p&gt;

&lt;p&gt;Yet the feedback we received from (beta) users on the 2022 beta versions of GermaParl v2 confirmed a suspicion: The data structure of GermaParl v2 was not yet sufficiently intuitive. So to improve the usability of the data, we somewhat reworked the data structure. The new beta release (GermaParl v2.0.0-beta.3) now available at &lt;a href=&quot;https://doi.org/10.5281/zenodo.7783309&quot;&gt;Zenodo&lt;/a&gt; addresses usability issues. This release note explains how the data structure has changed and how to use it.&lt;/p&gt;

&lt;h2 id=&quot;getting-started-with-germaparl-v200-beta3&quot;&gt;Getting started with GermaParl v2.0.0-beta.3&lt;/h2&gt;

&lt;p&gt;The &lt;a href=&quot;https://doi.org/10.5281/zenodo.7783309&quot;&gt;landing page of GermaParl v2.0.0-beta.3 at Zenodo&lt;/a&gt; explains how to register as a beta user and how to install the corpus using the ‘&lt;a href=&quot;https://CRAN.R-project.org/package=cwbtools&quot;&gt;cwbtools&lt;/a&gt;’ package.&lt;/p&gt;

&lt;p&gt;Using the more hierarchical data structure of GermaParl v2 requires changes of the packages used for analysing the data. In the background, the &lt;a href=&quot;https://CRAN.R-project.org/package=RcppCWB&quot;&gt;RcppCWB&lt;/a&gt; package has been developed stepwise to meet these requirements. Using the reworked data structure of GermaParl v2 requires the most recent release of the &lt;a href=&quot;https://CRAN.R-project.org/package=polmineR&quot;&gt;polmineR&lt;/a&gt; package: Installing the latest version (v0.8.8), recently published at CRAN is a prerequisite to leverage the potential of the new beta release. Install it if necessary, and load it.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;packageVersion&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;polmineR&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;numeric_version&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;0.8.8&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;install.packages&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;polmineR&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;polmineR&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;h2 id=&quot;so-this-is-new&quot;&gt;So this is new&lt;/h2&gt;

&lt;p&gt;The differences between the current and the previous beta release of GermaParl v2 are somewhat technical. Contrasting GermaParl v1 and GermaParl v2 may be the best entry point to explain what is new: GermaParl v2 moves beyond the deliberately “flat” data structure of GermaParl v1: The XML indexed as a corpus using the Corpus Workbench (CWB) functionality now maintains the logic that plenary protocols are the basic unit of data preparation. Information at the level of plenary protocols (date, legislative period, session number) is now maintained at this level, and is not turned into speech/speaker-level attributes as previously. The data structure with a hierarchy of plenary protocols, speakers, paragraphs, sentences and named entities (spans of words within sentences) can be visualized as follows.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/2023-04-03-GermaParl-v2-beta3-Release-Note/uml-1.png&quot; alt=&quot;plot of chunk uml&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The hierarchy of the data structure was somewhat concealed by GermaParl v1 and the previous beta release of GermaParl v2. The naming of structural attributes was kept as short as possible, removing information about the structure of the data. The recent GermaParl v2 release maintains this information and makes it explicit: The names of structural attributes now combine the XML element name and XML attribute names, pasted with an underscore. The following mapping table shows, how structural attributes of GermaParl v1 now correspond to structural attributes of GermaParl v2.&lt;/p&gt;

&lt;table&gt;
&lt;caption&gt;Mapping table for structural attributes (GermaParl/GermaParl2)&lt;/caption&gt;
 &lt;thead&gt;
  &lt;tr&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; GermaParl v1.0.6 &lt;/th&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; GermaParl v2.0.0-beta3 &lt;/th&gt;
  &lt;/tr&gt;
 &lt;/thead&gt;
&lt;tbody&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; lp &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_lp &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; session &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_no &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; date &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_date &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; year &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_year &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; src &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_filetype &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; url &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_url &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; agenda_item &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; agenda_item_type &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_name &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; party &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_party &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; parliamentary_group &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_parlgroup &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; role &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_role &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; interjection &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; p_type &lt;/td&gt;
  &lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;What is more: As we learned from users’ feedback, the previous beta version of GermaParl v2 maintained the “original” data structure of plenary protocols concerning interjections, but made it hard to handle. Interjections were XML children of speeches - making it difficult to exclude them from the analysis of speeches. But this is obviously necessary, if you want to analyse speech-making. Our practical solution now is to model interjections and stage instructions as distinct types of paragraphs within speeches. The introduction of the structural attribute &lt;em&gt;p_type&lt;/em&gt; is the precondition for subsetting subcorpora of speeches to exclude interjections.&lt;/p&gt;

&lt;p&gt;The following table conveys the full picture how the naming of s-attributes has changed between GermaParl v2.0.0-beta.2 and GermaParl v2.0.0-beta.3.&lt;/p&gt;

&lt;table&gt;
&lt;caption&gt;Mapping table for structural attributes (GermaParl v2: beta.2/beta.3)&lt;/caption&gt;
 &lt;thead&gt;
  &lt;tr&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; GermaParl v2.0.0-beta.2 &lt;/th&gt;
   &lt;th style=&quot;text-align:left;&quot;&gt; GermaParl v2.0.0-beta.3 &lt;/th&gt;
  &lt;/tr&gt;
 &lt;/thead&gt;
&lt;tbody&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; plenary_protocol &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; lp &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_lp &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_no &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_no &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; date &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_date &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; year &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_year &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; url &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_url &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; filetype &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; protocol_filetype &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_node &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; who &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_who &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_name &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; parpiamentary_group &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_parlgroup &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; party &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_party &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; role &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; speaker_role &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; stage &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; [not available] &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; stage_type &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; [not available] &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; p &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; p &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; [not available] &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; p_type &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; s &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; s &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; ner &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; ne &lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; ner_type &lt;/td&gt;
   &lt;td style=&quot;text-align:left;padding: 2px&quot;&gt; ne_type &lt;/td&gt;
  &lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The information, whether tokens are part of a speech or of an interjection reported in stage instructions is now part of the annotation of paragraphs. The structural attribute “p_type”, which assumes the values “speech” and “stage”, can now be used to subset a corpus to keep speeches and to drop interjections. See the following sample code.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;by_name&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;GERMAPARL2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;subset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;p_type&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;speech&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;subset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;speaker_name&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Angela Merkel&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;The polmineR package has been extended and reworked to be able to process the nested XML structure. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;subset()&lt;/code&gt; method for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;corpus&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;subcorpus&lt;/code&gt; objects has been introduced to offer much more flexibility for working with nested data structures. The most important difference with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;subset()&lt;/code&gt; for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;corpus&lt;/code&gt;/&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;subcorpus&lt;/code&gt; objects is that non-standard evaluation can be used in a way widely used in the tidyverse. Consider the following example how this is useful.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;by_date&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;GERMAPARL2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;subset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;p_type&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;speech&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;subset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;as.Date&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;protocol_date&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;as.Date&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;2001-09-11&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;One final point: The ability to display the full text of parliamentary speeches is a unique feature of polmineR. As things stand, it is necessary feed a paragraph-based (s-attribute p) definition of a subcorpus into &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;read()&lt;/code&gt; to be able to display the full text with formatted interjections - see the following example.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;speeches&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;GERMAPARL2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;subset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;protocol_date&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;2001-09-12&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;subset&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;as.speeches&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;s_attribute_date&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;protocol_date&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;s_attribute_name&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;speaker_name&quot;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;speeches&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[[&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;read&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;h2 id=&quot;where-to-go-from-here&quot;&gt;Where to go from here&lt;/h2&gt;

&lt;p&gt;This release note has a focus on the changes of the data structure of the new beta release of GermaParl v2 and how to use it with the polmineR package. Certainly, many users will consider the historical breadth of the data to be the most interesting aspect of GermaParl v2. But if the data is not modeled adequately and usable at the same time, analytical usage will be inhibited. As GermaParl v2.0.0-beta.3 is now available, we will be happy to receive further feedback that we can address by either improving the data, or the documentation.&lt;/p&gt;

&lt;p&gt;Our further release plan for GermaParl v2 now is as follows:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;A final beta release (#4) of GermaParl v2 is scheduled for late April 2023. This upcoming (last) beta version will improve data quality and offer extended documentation.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;We plan to officially release GermaParl v2 on this year’s Constitution Day, on May 23. Data structure, data quality and documentation shall then be consolidated to meet the requirements of a wider audience. The data will then be available without the need to register as a beta user.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The release of GermaParl v2 in May will be the starting point of further quality improvements based on user feedback. But full disclosure: In the context of our contribution to the &lt;a href=&quot;https://www.text-plus.org&quot;&gt;Text+&lt;/a&gt; consortium as part of &lt;a href=&quot;https://www.nfdi.de/nfdi4microbiota/?lang=en&quot;&gt;Germany’s Research Data Infrastructure (NFDI)&lt;/a&gt;, we are already working on an XML version of GermaParl that will then be based on the &lt;a href=&quot;https://www.clarin.eu/parlamint&quot;&gt;ParlaMint&lt;/a&gt; standard (upcoming GermaParl v3). And in the Kontext of &lt;a href=&quot;https://www.konsortswd.de/en/&quot;&gt;KonsortSWD&lt;/a&gt;, which is also part of NFDI, we will use GermaParl as public sample data for entity linking. So GermaParl v2 is an intermediate step - but an important one.&lt;/p&gt;
</description>
        <pubDate>Mon, 03 Apr 2023 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2023/04/03/GermaParl-v2-beta3-Release-Note.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2023/04/03/GermaParl-v2-beta3-Release-Note.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>The case for biglda: Topic modelling benchmarks</title>
        <description>&lt;h1 id=&quot;the-case-for-biglda-topic-modelling-benchmarks&quot;&gt;The case for biglda: Topic modelling benchmarks&lt;/h1&gt;

&lt;p&gt;As I picked up developing the biglda R package again (a GitHub-only package, see
&lt;a href=&quot;https://github.com/PolMine/biglda/tree/dev&quot;&gt;dev branch here&lt;/a&gt;), I started to
wonder: Is this is really worth the effort? Pleasure and pain are mixed, trying
to interface to Java via rJava for using
&lt;a href=&quot;https://mimno.github.io/Mallet/index&quot;&gt;Mallet&lt;/a&gt; as a topic modelling tool, and
when interfacing to Python via reticulate for using
&lt;a href=&quot;https://radimrehurek.com/gensim/&quot;&gt;Gensim&lt;/a&gt;, which is considered the top choice
for topic modelling in the Python realm.&lt;/p&gt;

&lt;p&gt;The evidence-based approach is to run a benchmark. I want to report results here
(including code) that would have bloated the &lt;a href=&quot;https://github.com/PolMine/biglda/tree/dev&quot;&gt;README of the biglda
package&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The purpose of the biglda package is to offer functionality and workflows to
cope with fitting topic models for large data set sets, i.e. to address performance
issues and issues with memory limitations that can be quite a nuisance with
biggish data. The package offers an interface to Gensim (Python) and Mallet
(Java), but the focus is on Mallet. Considering the results, I believe Mallet is
a very good choice for the scenarios I am currently dealing with - fitting
various topic models on a corpus with ~ 1 billion words and 2 million documents
to evaluate which k is a good choice.&lt;/p&gt;

&lt;p&gt;This does not mean that Mallet is my recommendation for any scenario -
absolutely not. Please do not skip the discussion at the end of this blog entry!
But for know, let me walk you through the benchmarking exercise. Along the way,
the code will also acquaint you (superficially) with the biglda package.&lt;/p&gt;

&lt;p&gt;The alternatives evaluated in this benchmark are:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Mallet as exposed to R via biglda&lt;/li&gt;
  &lt;li&gt;Gensim as exposed to R via biglda&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stm()&lt;/code&gt; of the &lt;a href=&quot;https://www.structuraltopicmodel.com&quot;&gt;stm&lt;/a&gt; R package (structural topic model)&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;LDA()&lt;/code&gt;, the classic of the &lt;a href=&quot;https://CRAN.R-project.org/package=topicmodels&quot;&gt;topicmodels&lt;/a&gt; R package&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a matter of convenience and reproducibility, we use the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AssociatedPress&lt;/code&gt;
data included in the topicmodels package, a representation of corpus data as a
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;DocumentTermMatrix&lt;/code&gt;, a class defined in the
&lt;a href=&quot;https://CRAN.R-project.org/package=tm&quot;&gt;tm&lt;/a&gt; package. The beauty of this data
format is that every tool considered is able to digest it. We also load the
‘slam’ package for functionality to process the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;DocumentTermMatrix&lt;/code&gt;.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;slam&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;data&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;AssociatedPress&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;package&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;topicmodels&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;Admittedly, the Associated Press (AP) data with 2246
documents and 10473 terms is not at all big data. But it is
large enough to get a good sense of the performance of different topic modelling
tools.&lt;/p&gt;

&lt;p&gt;Fitting a model with 100 topics is a good choice for the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AssociatedPress&lt;/code&gt;
corpus, see the respective evaluation in the &lt;a href=&quot;https://cran.r-project.org/web/packages/ldatuning/vignettes/topics.html&quot;&gt;vignette of the ldatuning
package&lt;/a&gt;.
Here, we just run 100 iterations, which will be enough to assess the performance
of the different implementations of topic modelling. In real life, you would run
more iterations. Mallet and Gensim are multi-threaded and can use multiple
cores. We use all but two cores (on a MacBookPro with 8 cores and an M1 Pro main
processor).&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;100L&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iterations&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;100L&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cores&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;parallel&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;::&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;detectCores&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;2L&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;As we run the tools, we fill a list successively.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;benchmarks&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;list&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;We start with Mallet, as exposed to R via biglda.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;biglda&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot; data-lang=&quot;text&quot;&gt;## Mallet version: v202108&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot; data-lang=&quot;text&quot;&gt;## JVM memory allocated: 0.5 Gb&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;instance_list&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;as.instance_list&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AssociatedPress&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;verbose&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;mallet_started&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BigTopicModel&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;n_topics&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;alpha_sum&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;5.1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;beta&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;0.1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;addInstances&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;instance_list&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;setNumThreads&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cores&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;setNumIterations&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iterations&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;setTopicDisplay&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;0L&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;0L&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# no intermediate report on topics&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;logger&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;setLevel&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rJava&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;::&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;J&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;java.util.logging.Level&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;OFF&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# remain silent&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BTM&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;estimate&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;benchmarks&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mallet&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;data.frame&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tool&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sprintf&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mallet_%d&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cores&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;as.numeric&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;difftime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;mallet_started&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;units&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mins&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;We continue with Gensim, which is run in a virtual Python environment from R
using &lt;a href=&quot;https://rstudio.github.io/reticulate/&quot;&gt;reticulate&lt;/a&gt;.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reticulate&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;gensim&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reticulate&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;::&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;import&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;gensim&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;bow&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dtm_as_bow&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AssociatedPress&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dict&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dtm_as_dictionary&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AssociatedPress&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;gensim_started&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;gensim_model&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;gensim&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;models&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ldamulticore&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;LdaMulticore&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;py&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;id2word&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;py&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dictionary&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;num_topics&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iterations&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iterations&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;per_word_topics&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;workers&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cores&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;benchmarks&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;gensim&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;data.frame&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tool&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sprintf&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;gensim_%d&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cores&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;as.numeric&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;difftime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;gensim_started&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;units&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mins&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;We then want look at the classic implementation of the
&lt;a href=&quot;https://CRAN.R-project.org/package=topicmodels&quot;&gt;topicmodels&lt;/a&gt; R package.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;topicmodels&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;topicmodels_started&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;lda_model&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;LDA&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AssociatedPress&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;method&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Gibbs&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;control&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;list&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iter&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iterations&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;benchmarks&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;topicmodels&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;data.frame&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tool&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;topicmodels&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;difftime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;topicmodels_started&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;units&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mins&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;Finally, we look at the structural topic model, the
&lt;a href=&quot;https://CRAN.R-project.org/package=stm&quot;&gt;stm&lt;/a&gt; package.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stm&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot; data-lang=&quot;text&quot;&gt;## stm v1.3.6 successfully loaded. See ?stm for help. 
##  Papers, resources, and other materials at structuraltopicmodel.com&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;AP&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;readCorpus&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AssociatedPress&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;type&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;slam&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stm_started&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stm_model&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stm&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;documents&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AP&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;documents&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;vocab&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;AP&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;vocab&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;K&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;k&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reportevery&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;0L&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;verbose&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;max.em.its&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;iterations&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;init.type&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;LDA&quot;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;benchmarks&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;stm&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;data.frame&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tool&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;stm&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;difftime&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Sys.time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stm_started&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;units&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;mins&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;And this is the result!&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-r&quot; data-lang=&quot;r&quot;&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ggplot2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ggplot&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;data&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;do.call&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rbind&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;benchmarks&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;aes&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tool&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;geom_bar&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stat&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;identity&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;&lt;img src=&quot;/assets/2023-02-15-Topic-Modelling-Benchmarks/benchmark_plot-1.png&quot; alt=&quot;plot of chunk benchmark_plot&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Quite clearly, Mallet is the fastest option. Good news for me! The effort I have
invested in developing the biglda package is not in vain. In fact, there is
another R package simply called mallet that exposes Mallet to R. Unfortunately,
it does not expose the multi-threaded &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ParallelTopicModel&lt;/code&gt;. This is what biglda 
does - plus some extra functionality to manage issues of memory limitations and 
efficiency at the R-Java-interface. So I am glad to find that biglda speeds up
topic modelling quite a bit.&lt;/p&gt;

&lt;p&gt;I am somewhat surprised to see that Gensim is so much slower, and even slower
than the topicmodels LDA implementation. I somehow suspect that running Gensim
in a virtual environment via reticulate may limit the access of Python to full
system resourses. Advice how the performance of Gensim could be improved is
welcome! And it is important to note that Gensim may deal with big data
exceeding the main memory of a machine (unlike Mallet), and Gensim is able to
use distributed computing. So even if Gensim is slower than Mallet, there are
big data scenarios with a strong case for Gensim.&lt;/p&gt;

&lt;p&gt;The comparatively weak performance of stm is the second surprise for me. But
there are many things that stm offers that are absolutely nice and beyond the
capabilities of Mallet (Gensim and topicmodels): Rich tools for evaluating topic
models and, of course, the capability of the algorithm to process metadata. So
stm is the perfect choice for a considerable bandwith of scenarios.&lt;/p&gt;

&lt;p&gt;So the message is: For fitting topic models for bulky data, if you need (various)
topic models within reasonable time and if memory limitations are an issue, 
Mallet is a great choice. And biglda is designed for R users to take advantage 
of the full potential of this topic modelling tool.&lt;/p&gt;

&lt;p&gt;Feedback welcome!&lt;/p&gt;
</description>
        <pubDate>Wed, 15 Feb 2023 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2023/02/15/Topic-Modelling-Benchmarks.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2023/02/15/Topic-Modelling-Benchmarks.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>Rcppcwb V0.4.4 Released</title>
        <description>&lt;p&gt;A new RcppCWB version “Jaberwocky” (v0.4.4) just made it to CRAN. Initially, this release was meant to be a minor maintenance release to address a warning on paths in an example of the &lt;a href=&quot;https://CRAN.R-project.org/package=cwbtools&quot;&gt;cwbtools&lt;/a&gt; package. The exercise went beyond that, a broader set of issues and bug reports have been addressed, to make RcppCWB an efficient, robust and trustworthy basis for processing text as linguistic data.&lt;/p&gt;

&lt;p&gt;RcppCWB has become much more consistent in handling paths - a potential source of confusing messages and errors. The basic issue is that the &lt;a href=&quot;https://cwb.sourceforge.io/&quot;&gt;Corpus Workbench (CWB)&lt;/a&gt; is parses the registry files describing corpora only once. After loading a corpus, an internal C representation keeps information on a corpus. But when the corpus is modified in any way (i.e. by adding an s-Attribute with a new annotation layer), this internal C represenation may be outdated. RcppCWB previously did not consider this siutation appropriately and only offered functionality prone to crash.&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cl_deleted_corpus()&lt;/code&gt; function to unload a corpus is now robust.  A new utility function &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;corpus_is_loaded()&lt;/code&gt; offers a check. Various functions have been updated to handle paths more consistently across platforms (Linux, macOS and Windows), to avoid confusing error messages.&lt;/p&gt;

&lt;p&gt;The release also includes an extended test suite and is an intermediate step to achieve full Windows compatibility of the functionality to build corpora. This is still to be achieved for the functions &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cwb_encode()&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cwb_makeall()&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cwb_compress_rdx()&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;cwb_huffcode()&lt;/code&gt;). Note that this is a limitation to build corpora using RcppCWB on Windows. There are no known limitations of the functionality to analyzing corpus data.&lt;/p&gt;

&lt;p&gt;Further work on Windows compatibility is a milestone envisaged for RcppCWB v0.6.0. For RcppCWB v0.5.0, the goal is to re-align the CWB code with the latest version of RcppCWB.&lt;/p&gt;
</description>
        <pubDate>Tue, 14 Dec 2021 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/2021/12/14/RcppCWB-v0.4.4-released.html</link>
        <guid isPermaLink="true">http://polmine.github.io/2021/12/14/RcppCWB-v0.4.4-released.html</guid>
        
        
      </item>
    
      <item>
        <title>cwbtools v0.3.3 &apos;Hemicycle&apos; brings Europarl closer. </title>
        <description>&lt;h1 id=&quot;cwbtools-v033-hemicycle-brings-europarl-closer&quot;&gt;cwbtools v0.3.3 ‘Hemicycle’ brings Europarl closer.&lt;/h1&gt;

&lt;p&gt;As an immediate follow-up to cwbtools v0.3.2 “Il Postino”, a new cwbtools version (v0.3.3, “Hemicycle”) just made it to CRAN. This has become necessary to fix errors with the Solaris and Fedora test environments of CRAN that occurred because v0.3.2 expanded test coverage.&lt;/p&gt;

&lt;p&gt;Changes to meet CRAN requirements will not be relevant for most users. So why bother to install the new release? There is one nice new feature: Assumptions made by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;corpus_install()&lt;/code&gt;, the mechanism for installing corpora, have been relaxed as follows: The binary data directory in a corpus tarball may now also be named “data” (not only “indexed_corpora”) and binary files need not reside in a subdirectory named after the corpus. They can be in the data directory directly.&lt;/p&gt;

&lt;p&gt;Access to two classic corpora - the Europarl corpus with debates in the European parliament and the Dickens corpus with Charles Dickens’ novels - is super-easy now: Simply use the following code to download and install both corpora.&lt;/p&gt;

&lt;div class=&quot;language-r highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cwbtoos&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dickens&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;http://cwb.sourceforge.net/temp/Dickens-1.0.tar.gz&quot;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus_install&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tarball&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dickens&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ep&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;http://corpora.linguistik.uni-erlangen.de/demos/download/Europarl3-CWB-2010-02-28.tar.gz&quot;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;corpus_install&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tarball&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ep&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;A second note on this release: The functionality that has evolved to install corpora from Web storage, from a repository such as Zenodo and from Amazon S3 supersedes previous ideas to ship (large) corpora in R data packages. The big disadvantage of using R data packages for data dissemination is the need to re-install corpora every time there is a major R update. So cwbtools v0.3.3 makes initial steps to declare functionality as deprecated that assists users to wrap corpora in packages. In this vein, the previous &lt;em&gt;Europarl&lt;/em&gt;-vignette that explained how to wrap the Europarl corpus into an R data package has been dropped.&lt;/p&gt;

&lt;p&gt;But functionality to download and install Europarl from the Web is there! A nice thing about Europarl is that it is a multilingual ressource, making it relevant for many international users. As a tribute to the European Parliament (EP), this release is labelled “Hemicycle”: This is how the plenary chamber of the EP at Strasbourg is called, following a tradition to arrange seats and parliamentary groups in a circular shape.&lt;/p&gt;

</description>
        <pubDate>Tue, 23 Feb 2021 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2021/02/23/cwbtools-v0.3.3-Hemicycle-install-Europarl-at-ease.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2021/02/23/cwbtools-v0.3.3-Hemicycle-install-Europarl-at-ease.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>cwbtools v0.3.2 &apos;Il Postino&apos; solidifies corpus download. </title>
        <description>&lt;h1 id=&quot;rcppcwb-v032-il-postino-solidifies-corpus-download&quot;&gt;RcppCWB v0.3.2 ‘Il Postino’ solidifies corpus download.&lt;/h1&gt;

&lt;p&gt;A new, previously unknown bug users reported when trying to use the &lt;a href=&quot;https://CRAN.R-project.org/package=GermaParl&quot;&gt;GermaParl&lt;/a&gt; package to download and install the GermaParl corpus (function &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;germaparl_install_corpus()&lt;/code&gt;) generated some urgency to release of a new version of the cwbtools R package: We suddenly saw errors (false positives!) when checking whether GermaParl can be downloaded from Zenodo. A switch from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;RCurl::url.exists()&lt;/code&gt; to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;httr::http_error()&lt;/code&gt; solves this small yet nasty problem.&lt;/p&gt;

&lt;p&gt;A second important bug fix is that it is now possible to encode more than one positional attribute at a time – without hacks. We are relieved that this limitation of the usefulness of cwbtools is overcome.&lt;/p&gt;

&lt;p&gt;But cwbtools v0.3.2 also brings new features:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;When a corpus is downloaded from Zenodo, the md5 checksum is checked to ensure the integrity of the data.&lt;/li&gt;
  &lt;li&gt;It is now possible to download corpora from Amazon S3. This may be very useful for projects that cannot yet publish data publicly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So data delivery is more robust and more comprehensive. This is why the release is labelled ‘Il Postino’. To learn more about cwbtools v0.3.2 ‘Il Postino’, please consult the &lt;a href=&quot;https://polmine.github.io/cwbtools/news/index.html&quot;&gt;package changelog&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Many thanks to the users who reported the issue that downloading GermaParl did not work, temporarily! Good news: Downloading GermaParl from Zenodo works again the way it should.&lt;/p&gt;
</description>
        <pubDate>Wed, 17 Feb 2021 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2021/02/17/cwbtools-v0.3.2-Il-Postino-solidifies-and-expands-corpus-download.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2021/02/17/cwbtools-v0.3.2-Il-Postino-solidifies-and-expands-corpus-download.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>RcppCWB v0.3.2 &apos;Dune Ride&apos; ensures Apple Silicon compatibility. </title>
        <description>&lt;h1 id=&quot;rcppcwb-v032-dune-ride-ensures-apple-silicon-compatibility&quot;&gt;RcppCWB v0.3.2 ‘Dune Ride’ ensures Apple Silicon compatibility.&lt;/h1&gt;

&lt;p&gt;Apple’s new M1 chip has deservedly raised significant attention. The new Mac minis and 13 inch Macbook Pros running on “Apple Silicon” are fast, energy-saving and affordable at the same time. No doubt an increasing number of polmineR users will work on an M1 machine. So we need to take care that everything works with the new generation of processors.&lt;/p&gt;

&lt;p&gt;Technically, everything hinges on the &lt;a href=&quot;https://polmine.github.io/RcppCWB/&quot;&gt;RcppCWB&lt;/a&gt; package. The C code of the &lt;a href=&quot;http://cwb.sourceforge.net/&quot;&gt;Corpus Workbench (CWB)&lt;/a&gt; included in the RcppCWB package and the C++ wrappers exposing CWB performance are the basis for everything that makes analysing corpora with polmineR fast. So the crucial question was whether RcppCWB would compile on M1 machines?&lt;/p&gt;

&lt;p&gt;If you run R with the &lt;a href=&quot;https://en.wikipedia.org/wiki/Rosetta_(software)&quot;&gt;Rosetta&lt;/a&gt; emulator, everything is fine. But one issue arising with native M1 compiled code required our attention: If &lt;a href=&quot;https://developer.gnome.org/glib/&quot;&gt;GLib-2.0&lt;/a&gt;, a dependency of the CWB, is not yet present, it would be downloaded from the &lt;a href=&quot;https://github.com/PolMine/libglib/&quot;&gt;libglib&lt;/a&gt; repository we host with GitHub. And GLib binaries for the new processor architecture were not there yet. And RcppCWB, which had been developed assuming that every Mac has Intel inside, would not look for M1-compatible dependencies either.&lt;/p&gt;

&lt;p&gt;This is solved with &lt;a href=&quot;https://CRAN.R-project.org/package=RcppCWB&quot;&gt;RcppCWB v0.3.2&lt;/a&gt; which has just made its way to CRAN. The functionality of the package remains unchanged. Working on the configure script anyway, I tried to make it more robust and verbose so that users could get a better idea what is wrong when facing an issue when compiling the package. For instance, there is now a warning when &lt;a href=&quot;https://www.pcre.org/&quot;&gt;PCRE&lt;/a&gt; is present but has been compiled without Unicode support. Yet the primary focus of the new release is to ensure that RcppCWB (and polmineR, cwbtools etc) work seamlessly on the new Apple machines with M1 chips inside.&lt;/p&gt;

&lt;p&gt;This has been achieved. Because it is just not possible to get one of the new machines really quickly (delivery takes a few weeks), we used a cloud solution offered by MacStadium to be able to adjust things and run tests now. This worked nicely, no complaints. Nevertheless, making the adjustments required some attention. This is why this release of RcppCWB is called “Dune Ride”.&lt;/p&gt;

&lt;p&gt;You need further explanation? Dunes are made of sand, and processors are made of silicon, and silicon is made of sand, and so is the M1 chip, I guess, so rodeo with M1 is a dune ride, metaphorically.&lt;/p&gt;

&lt;p&gt;If you face issues running RcppCWB/polmineR/cwbtools on M1, please use &lt;a href=&quot;https://github.com/PolMine/RcppCWB/issues&quot;&gt;RcppCWB issues&lt;/a&gt; to let us know!&lt;/p&gt;
</description>
        <pubDate>Thu, 04 Feb 2021 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2021/02/04/RcppCWB-v0.3.2-Dune-Ride-Ensures-Apple-Silicon-Compatibility.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2021/02/04/RcppCWB-v0.3.2-Dune-Ride-Ensures-Apple-Silicon-Compatibility.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
      <item>
        <title>New Project &apos;Linking Textual Data&apos; started. </title>
        <description>&lt;h1 id=&quot;new-project-linking-textual-data-started&quot;&gt;New Project ‘Linking Textual Data’ started.&lt;/h1&gt;

&lt;p&gt;Our newest project “Linking Textual Data” is part of the &lt;em&gt;Consortium for the Social, Behavioral, Educational and Economic Sciences&lt;/em&gt; (&lt;a href=&quot;https://www.konsortswd.de/en/&quot;&gt;KonsortSWD&lt;/a&gt;) and contributes to the &lt;em&gt;National Research Data Infrastructure&lt;/em&gt; (&lt;a href=&quot;https://www.nfdi.de/en-gb&quot;&gt;NFDI&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;In this new project that has officially started on Januar 1st 2021, we develop tools for linking different data types in order to open up new avenues for research. Corpora have become an important data type in the social sciences. Methods for analysing text are developing rapidly. A crucial aspect of the huge potential of large-scale corpora for the social sciences are capabilities to link the analysis of text and other data types, such as surveys. Yet the barriers to data linkage are still high. The tools and workflows we will develop and share shall improve our abilities to gain new insights through the combination of different data types. Not surprisingly, corpora play a central role in our plans.&lt;/p&gt;

&lt;p&gt;The NFDI is the broader context of our endeavour. It aims to secure and utilize research data in a systematic and sustainable way. To reach this goal, the NFDI plans to establish the management of research data according to the “FAIR” principles. These principles mandate that research data are findable, accessible, interoperable and reusable for and by the scientific community. Furthermore, the NFDI aims to connect to international initiatives with similar aims such as the &lt;em&gt;European Open Science Cloud&lt;/em&gt; (&lt;a href=&quot;https://eosc-portal.eu&quot;&gt;EOSC&lt;/a&gt;). The efforts of the NFDI are not limited to particular disciplines, they include consortia as diverse as chemistry, culture, health, engineering and many more.&lt;/p&gt;

&lt;p&gt;Within the NFDI, the social sciences are represented through KonsortSWD. Specifically, KonsortSWD strengthens, widens and deepens a research data infrastructure for the social, educational, behavioural and economic sciences. Again, the FAIR principles offer guidance: KonsortSWD wants to provide the community (including research data centres) with adequate tools to share and manage data accordingly. Apart from community engagement, data access, ethics and technical solutions, another important task concerns the production of data. This is where our project is located. More specifically, our aim is to contribute to new insights by developing tools for data linkage.&lt;/p&gt;

&lt;p&gt;All efforts to utilize research data need to serve the community! Are there any best practices, opportunities and limits? What does the community need from such an infrastructure to be able to use it and profit from it? Community involvement is thus the key to successful data management. If you have any ideas or input how your projects could profit from “Linking Textual Data”, feel free to reach out to us.&lt;/p&gt;

</description>
        <pubDate>Wed, 06 Jan 2021 00:00:00 +0000</pubDate>
        <link>http://polmine.github.io/posts/2021/01/06/Linking-Textual-Data.html</link>
        <guid isPermaLink="true">http://polmine.github.io/posts/2021/01/06/Linking-Textual-Data.html</guid>
        
        <category>news</category>
        
        
        <category>Posts</category>
        
      </item>
    
  </channel>
</rss>
