Friday, March 17, 2017

List of EU-funded projects for language and translation

A | B | C | D | E | F | G | H | I | J | K | L | M | N | O | P | Q | R | S | T | U | V | W | X | Y | Z

Complete list of projects

Please note that the project factsheets will no longer be updated. All information relevant to the projects can be found on the CORDIS factsheets. These are updated on a regular basis with public deliverables, etc. Links to the CORDIS factsheets are given below, and on the respective project pages.

A

Back to top

accept-logo.png
ACCEPT - Automated Community Content Editing PorTal
The use of machine translation (MT) is becoming much more pervasive. At the same time, Web 2.0 paradigms are democratising content creation - stressing the value of communities of users creating content for each other. However, right now these two trends are fairly incompatible. MT engines, even statistical engines, cannot produce acceptable results for community content due to the extreme variability within the content. The ACCEPT project will address this issue by developing new technologies designed specifically to help MT work better in this environment.
Start date: 1 January 2012
End date:
31 December 2014
Project officer: Pierre-Paul Sondag
Factsheet on Cordis
Factsheet
Website
accurat-logo.jpg
ACCURAT - Analysis and Evaluation of Comparable Corpora for Under Resourced Areas of Machine Translation
Lack of sufficient linguistic resources for many languages and domains currently is one of the major obstacle in further advancement of automated translation. The main goal of the ACCURAT research is to find, analyze and evaluate novel methods how comparable corpora can compensate for this shortage of linguistic resources to improve MT quality significantly for under-resourced languages and narrow domains. The ACCURAT project will provide researchers and developers with novel methodology and fully functional model for exploiting comparable corpora to increase translation quality of existing and emerging MT systems.
Start date: 1 January 2010
End date:
30 June 2012
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Factsheet
Website
annomarket-logo-rectangular.png
AnnoMarket - Annotation Resource Marketplace in the Cloud
The project aims to revolutionise the text annotation domain, by delivering an affordable, open marketplace for pay-as-you-go, cloud-based extraction resources and services, in multiple languages. This project will be driven by a consortium dominated by commercial partners, from three EU countries and with 43% of the budget assigned to SMEs.
Start date: 1 June 2012
End date: 31 May 2014
Project officer: Susan Fraser
Factsheet on CORDIS
Website
atlas-logo.png
ATLAS - Applied Techology for Language-Aided CMS
The advent of the Web revolutionized the way in which content is manipulated and delivered. As a result, digital content in various languages has become widely available on the Internet and its sheer volume and language diversity have presented an opportunity for embracing new methods and tools for content creation and distribution. Although significant improvements have been made lately in the field of web content management, there is still a growing demand for online content services that incorporate language-based technology. Mechanisms such as automatic annotation of important words, phrases and names, text summarization and categorization, and computer-aided translation could facilitate the process of manipulating heterogeneous multilingual content as well as enhance end-user experience by allowing for better content navigation. This project unifies such mechanisms in a common software platform called ATLAS and builds three separate solutions around this platform.
Start date: 1 March 2010
End date:
28 February 2013
Project officer: Pierre-Paul Sondag
Factsheet on EUROPA
Factsheet
Website

B

Back to top

bologna-logo.png
BOLOGNA - Bologna Translation Service
There is a continuing increasing need for educational institutes to provide course syllabi documentation and other educational information in English. Access to translated course syllabi and degree programmes plays a crucial role in the degree to which universities effectively attract foreign students and, more importantly, have an impact on international profiling. To present all education information in English is a major challenge for most higher educational institutes and the figures and trends show that investment in traditional human translation services is prohibitive, consequently course materials and degree programmes are often provided in local languages only. BOLOGNA aims to provide a solution to the problem by offering a low-cost, web-based, high-quality translation service already in use by several Swedish universities.
Start date: 1 March 2011
End date:
28 February 2013
Project officer: Pierre-Paul Sondag
Factsheet on EUROPA
Factsheet
Website

C

Back to top

casmacat-logo.jpg
CASMACAT - Cognitive Analysis and Statistical Methods for Advanced Computer Aided Translation
The CASMACAT project will build the next generation translators workbench to improve productivity, quality, and work practices in the translation industry. We will carry out cognitive studies of actual unaltered translator behaviour based on key logging and eye tracking. The acquired data will be examined for how interfaces with enriched information are used, to determine translator types and styles, and to build a cognitive model of the translation process.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Factsheet
Website
cesar-logo.bmp
CESAR - Central and South-East European Resources
Human language technologies crucially depend on language resources and tools that are useable, useful and available. However, even where language resources and respective tools are available they have been developed mostly in a sporadic manner, in response to specific project needs and with relatively little regard to their long-term sustainability, IPR status, interoperability, reusability in different contexts as well as to their potential deployment in multilingual applications. CESAR, in close harmony with META-NET intends to address this issue by enhancing, upgrading, standardising and cross-linking a wide variety of language resources and tools and making them available, thus contributing to an open linguistic infrastructure.
Start date: 1 February 2011
End date:
31 January 2013
Project officer: Kimmo Rossi
Factsheet on EUROPA
Factsheet
Website
CLASSiC_logo.jpg
CLASSiC - Computational Learning in Adapptive systems for Spoken Conversation
The overall goal of the CLASSiC project is to facilitate the rapid deployment of accurate and robust spoken dialogue systems that can learn from experience. The approach is based on statistical learning methods with a unified treatment of uncertainty across the entire systems (speech recognition, spoken language and understanding, dialogue management, natural language generation and speech synthesis). This will result in a modular processing framework with an explicit representation of uncertainty connecting the various sources of uncertainty (understanding errors, ambiguity, etc) to the constraints to be exploited (task, dialogue and user contexts). the architecture supports a layered hierarch of supervised learning and reinforcement learning in order to facilitate mathematically principled optimisation and adaptation techniques. It is being developed in close co-operation with an industrial partner in order to ensure a practical deployment platform as well as a flexible research test-bed.
Start date: 1 March 2008
End date: 28 February 2011
Project officer: Philippe Gelin
Factsheet on CORDIS
Website

Co-friend - Cognitive and flexible learning system operating robust interpretation of extended real scenes by multi-sensors datafusion
Co-FRIEND aims to design a framework for understanding human activities in real environments, through an artificial cognitive vision system, identifying objects and events, and extracting sense from scene observation. It will manage uncertainty and change, and will create analysis meaning.
Start date: 1 February 2008
End date: 28 February 2011
Project officer: Michel Brochard
Factsheet on CORDIS
cosyne-logo.jpg
COSYNE - Multi-Lingual Content Synchronisation with WIKIS
The combination of dynamic user-generated content and multi-lingual aspects is particularly prominent in Wiki sites. Wikis have gained increased popularity over the last few years as a means of collaborative content creation as they allow users to set up and edit web pages directly. A growing number of organizations use Wikis as an efficient means to provide and maintain information across several sites. Currently, multi-lingual Wikis rely on users to manually translate different Wiki pages on the same subject. This is not only a time-consuming procedure but also the source of many inconsistencies, as users update the different language versions separately, and every update would require translators to compare the different language versions and synchronize the updates.
Start date: 1 March 2010
End date:
28 February 2013
Project officer: Stefano Bertolo
Factsheet on CORDIS
Website

D

Back to top

dictasign logo
Dicta-Sign - Sign language Recognition, Generation and Modelling with application in Deaf Communication
Dicta-Sign addresses the need for communication between deaf individuals and communication via natural language by deaf users with various human computer interfaces (HCI) environments. Dicta-Sign is one of two FP7 projects addressing sign language. The ultimate goal is to enable deaf users to fully integrate into the information society and use interactive social media in their mother tongue, the sign language, and in communication between themselves and with hearing people who use "normal" written language. The technology to be developed will enable automatic recognition of sign language, transcription and translation into written language, and automatic generation of sign language by online avatars.
Start date: 1 February 2009
End date: 31 January 2012

Project officer: Kimmo Rossi
Factsheet on CORDIS
Factsheet
Website
dirha-logo.png
DIRHA - Distant speech Interaction for Robust Home Applications
The DIRHA project addresses the development of voice-enabled automated home environments based on distant-speech interaction in different languages. A distributed microphone network is installed in the rooms of a house in order to monitor selectively acoustic and speech activities observable inside any space, and to eventually run a spoken dialogue session with a given user in order to implement a service or to have access to appliances and other devices.
Start date: 1 January 2012
End date:
31 December 2014
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website

E

Back to top

eastin-cl-logo.png
EASTIN-CL - Crosslingual and multimodal Search in a Portal for Support of Assisted Living
The project will support the social participation of disabled and elderly people by providing a crosslingual and multimodal portal containing information on assistive tools and technology. There are large collections of assistive technology products and additional information, searched by both professionals (doctors, insurances, manufacturers) and end users. Eastin-CL will improve access to this information.
Start date: 1 March 2010
End date:
31 May 2012
Project officer: Michel Brochard
Factsheet on EUROPA
Factsheet
EMIME_logo.jpg
EMIME - Effective Multilingual Interaction in Mobile Environments
The EMIME project will help to overcome the language barrier by developing a mobile device that performs personalised speech-to-speech translation, such that the a user’s spoken input in one language is used to produce spoken output in another language, while continuing to sound like the user’s voice. Personalisation of systems for cross-lingual spoken communication is an important, but little explored, topic. It is essential for providing more natural interaction and making the computing device a less obtrusive element when assisting human-human interactions.
Start date: 1 March 2008
End date: 28 February 2011
Project officer: Philippe Gelin
Factsheet on CORDIS
Website
eumssi logo
EUMSSI - Event Understanding through Multimodal Social Stream Interpreting
The main objective of EUMSSI is developing technologies for identifying and aggregating data presented as unstructured information in sources of very different nature (video, image, audio, speech, text and social context), including both online (e.g., YouTube) and traditional media (e.g. audiovisual repositories), and for dealing with information of very different degrees of granularity. The multimodal analytics will help organize, classify and cluster cross-media streams, by enriching its associated metadata. A core idea is that the process of integrating content from different media sources is carried out in an interactive manner, so that the data resulting from one media helps reinforce the aggregation of information from other media, in a cross-modal interoperable semantic representation framework. This will be accomplished thanks to the integration in a multimodal platform of state-of-the-art information extraction and analysis techniques from the different fields involved. Interoperability and interactive reinforcement of the data aggregation and a high-level semantic, conceptual and eventive representation will distinguish this proposal from others that incorporate multimodal search. The resulting platform will be potentially useful for any application in need of cross-media data analysis and interpretation, such as intelligent content management systems, personalized recommendation, real time event tracking, content filtering, etc. The project brings together 5 universities and research centres, a public service broadcaster and a SME providing solutions for the media industry. The real-world necessities of the 2 user partners motivate two strong user cases that have immediate market applicability. We also expect EUMSSI, which covers English, German, Spanish and French, to promote interaction and mutual knowledge among the diverse linguistic communities within Europe.
Start date: 1 December 2013
End date:
30 November 2016
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Website
eubridge-logo.gif
EU-BRIDGE - Bridges Across the Language Divide
Today Europe is facing larger and more critical language challenges than ever before. The production of multilingual content now far outpaces our ability to translate it by human effort and we must turn to automatic methods to cope. Thus, effective and innovative alternatives must be provided to Europes citizens and businesses. High performing machine translation technology can be part of the solution. Recent advances in machine translation (MT) technology now show great promise, as systems can be trained automatically from data and achieve respectable performance, even from speech input. However, MT still has very high maintenance costs, and is unsuited to cope with many of todays digital medias relentlessly changing streams of information, across different topics, styles, and genres. "Bridges Across the Language Divide" (EU-BRIDGE) proposes to advance speech translation to the point where it can deal with the varying input conditions occurring in digital media, and is able to automatically adapt itself to the changing domains.
Start date: 1 February 2012
End date:
31 January 2015
Project officer: Carola Carstens
Factsheet on CORDIS
Factsheet
Website
euromatrixplus logo
EuroMatrixPlus - Bringing Machine Translation for European Languages to the User
EuroMatrixPlus is our "flagship" project in machine translation. It addresses all 23 EU official languages and progressively improves the quality of automatic translations by moving towards the self-learning machine translation (able to learn from its mistakes). To be noted that this theme is currently being expanded by recent projects from FP7 Call 4, currently under negotiation.
Europe provides a challenge due to its vast diversity of languages, and machine translation technology will provide means to address this challenge. EuroMatrixPlus will lift the research strategy and infrastructure for statistical and hybrid machine translation developed within EuroMatrix to a higher level by adding new scientific components, by experimenting with novel community-based methods for testing and data collection and by building and utilizing stronger connections into the user world. At the same time EuroMatrixPlus will preserve the established brand name for the well accepted infrastructure, continue the most promising research strands, and extend the growing open source research community hat has formed around the EuroMatrix project.
Start date: 1 March 2009
End date: 30 April 2012

Project officer: Michel Brochard
Factsheet on CORDIS
Factsheet
Website
eurosentiment-logo.png
EUROSENTIMENT - Language Resource Pool for Sentiment Analysis in European Languages
During the last years, there has been a high increase in the use of social networks and blogs so that citizens and consumers express now widely their opinions about different topics like politics, society and media, through these channels. However the development of systems for sentiment analysis of these opinions is hampered by difficulties to access and get the necessary language resources, for several reasons:
- language resource owners fears for losing competitiveness;
- lack of agreed language resource schemas for sentiment analysis and not normalised magnitudes for measuring sentiment strength;
- high costs for adapting existing language resources for sentiment analysis;
- reduced visibility, accessibility and interoperability of the language resources.

Start date: 1 September 2012
End date:
31 August 2014
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website
excitement-logo.jpg
EXCITEMENT - EXploring Customer Interactions Through textual EntailMENT
Identifying semantic inference relations between texts is a major underlying language processing task, needed in practically all text understanding applications. For example, Question Answering and Information Extraction systems should verify that extracted answers and relations are indeed inferred from the text passages; multi-document text summarization needs to infer that one sentence entails another in order to avoid redundantly including both in a summary; and so on. While such apparently similar inferences are broadly needed, there are currently no generic semantic "engines" or platforms for broad textual inference. Rather, tools exist for narrow semantic tasks, but systems have to independently assemble and augment them to obtain a complete inference process.
Start date: 1 January 2012
End date:
31 December 2014
Project officer: Carola Carstens
Factsheet on CORDIS
Factsheet
Website

F

Back to top

faust-logo.gif
FAUST - Feedback Analysis for User adaptive Statistical Translation
The FAUST project will develop machine translation (MT) systems which respond rapidly and intelligently to user feedback. Current web-based MT systems provide high-volume translation without real-time. Most systems provide no opportunity for users to offer opinions or corrections for translation results. Other systems ask users for feedback on translation, however the user does not see any benefit to providing feedback: the translation does not change in response to the feedback. Our goal is to develop high-volume translation systems capable of adapting to user feedback in real-time.
Start date: 1 February 2010
End date:
31 January 2013
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website
FLaReNet was submitted and selected for funding under the eContentplus Programme. This programme came to an end on 31 December 2008. Measures to make digital content more accessible will continue however under the ICT Policy Support Programme ( ICT-PSP ).
FLaReNet logo
FLaReNet
International cooperation and re-creation of a community are the most important drivers for a coherent evolution of the Language Resource (LR) area in the next years. FLaReNet will be a European forum to facilitate interaction among LR stakeholders. Its structure considers that LRs present various dimensions and must be approached from many perspectives: technical, but also organisational, economic, legal, political. The Network addresses also multicultural and multilingual aspects, essential when facing access and use of digital content in today's Europe.
The Proceedings of the FLaReNet Launch event "Shaping the Future of the Multilingual Digital Europe" held on 12-13 February 2009 in Vienna are available for download below.
FLaReNet Proceedings pdf.gif (1.17MB)
Start date: 1 September 2008
End date: 31 August 2011

Project officer: Kimmo Rossi
Factsheet on CORDIS
Website

FLAVIUS - Foreign LAnguage Versions of Internet and User generated Sites
The web has become the largest place to publish and share all kind of information, from news to user-generated content. The answer to almost any question can be found online, but not necessarily in the working language of each user. Thus, despite of the development of translation tools, language is still a barrier, as most people access information written in their own language only. The FLAVIUS project aims at bridging the language gap between content publishers and users by providing an online platform accessible to websites owners that will enable them to generate multilingual versions of their site, quickly, easily and efficiently in as many languages as they want.
Start date: 1 April 2010
End date:
30 September 2012
Project officer: Aleksandra Wesolowska
Factsheet on EUROPA
Factsheet
Website

G

Back to top

galateas-logo.gif
GALATEAS - Generalized Analysis of Logs for Automatic Translation and Episodic Analysis of Searches
With the growth of digital libraries and digital library federation (as well as partially unstructured collections of documents such as web sites), a large set of vendors is offering engines for retrieving contents and metadata via search requests by the end user (queries). In most cases these queries are just unstructured fragments of text in a specific language. The first service offered by GALATEAS (LangLog) is focussed on getting meaning out of these lists of queries and it is addressed to library/federation/site managers.
Start date: 1 April 2010
End date:
31 March 2013
Project officer: Pierre-Paul Sondag
Factsheet on EUROPA
Factsheet
Website
gethomesafe-logo.jpg
GET HOME SAFE - Extended Multimodal Search and Communication Systems for Safe In-Car Application
The aim of the proposed project is to develop a system for safe information access (search, navigation, point of interest) and communication (texting) while driving. In order to reach that goal, we approach the problem from a holistic view, investigating the underlying "driving forces", studying the goals underlying searching and texting. Special attention will be paid to task and context factors such as multi-tasking and the associated cognitive load for the driver.
Start date: 1 January 2012
End date:
31 December 2014
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website

I

Back to top

itranslate4-logo.png
iTRANSLATE4 - Internet Translators for all European Languages
The goal of this project is to integrate the best machine translation services of all the major European MT providers in a single website that will offer free online machine translation from any official European language to any other. Translation between all European language pairs will be available by the partners direct or linked translators. Service providers will advertise their extra translation services and also make profit from adverts sold here.
Start date: 1 March 2010
End date:
29 February 2012
Project officer: Philippe Gelin
Factsheet on EUROPA
Factsheet
Website

L

Back to top

letsmt.gif
LetsMT! - Platform for Online Sharing of Training Data and Building User Tailored MT
In recent years, statistical machine translation (SMT) has become the leading paradigm for machine translation. SMT systems are built by analyzing huge volumes of parallel corpus and learning translation models from this data. The quality of SMT systems largely depends on the size of training data. Since the majority of parallel data is in major languages, SMT systems for larger languages are of much better quality compared to systems for smaller languages. The cost and the know-how required for building custom MT solutions deter many small-to-medium companies from utilizing the power of MT technologies. To fully exploit the huge potential of existing open SMT technologies we propose to build an innovative online collaborative platform for data sharing and MT building.
Start date: 1 March 2010
End date:
31 August 2012
Project officer: Kimmo Rossi
Factsheet on EUROPA
Factsheet
Website
lider logo
LIDER - Linked Data as an enabler of cross-media and multilingual content analytics for enterprises across Europe
The explosive growth of content in volume, velocity and variety on the Web demands new approaches to content analytics, addressing issues in large scale analysis and interpretation of heterogeneous data sets, originating in different media, human languages, jurisdictions, etc. Among these, language diversity in particular has become a ubiquitous aspect of the Web in light of increasing globalization. Recently, semantic-level, i.e., language- and media-independent, data analysis and representation methods such as those provided by Linked Data and Semantic Web technologies, have been introduced to provide innovative content analytics solutions for such heterogeneous, multilingual and multimedia content. An important missing component however, is the representation of language- and media-specific information that will be needed for interpreting such data correctly - across different media and across the increasing variety of human languages used nowadays on the Web.
Start date: 1 November 2013
End date:
31 October 2015
Project officer: Susan Fraser
Factsheet on CORDIS
Website
limosine-logo.png
LiMoSINe - Linguistically Motivated Semantic aggregation engiNes
We increasingly live our life online. Information is accumulated on a wide range of human activities, from science and facts, to personal content, opinions, and trends. Across the globe, peoples knowledge, experiences and interactions effortlessly find their way to online outlets, alongside traditional edited content, ready to be shared with millions. LiMoSINe will integrate the research activities of leading researchers across diverse topics with a view to enabling new kinds of language-based search technology.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website
LIREC_logo.jpg
LIREC - Living with robots and interactive companions
LIREC aims to establish a multi-faceted theory of artificial long-term companions (including memory, emotions, cognition, communication, learning, etc.), embody this theory in robust and innovative technology and experimentally verify both the theory and technology in real social environments. Whether as robots, social toys or graphical and mobile synthetic characters, interactive and sociable technology is advancing rapidly. However, the social, psychological and cognitive foundations and consequences of such technological artefacts entering our daily lives - at work, or in the home - are less well understood.
The project has produced some entertaining and educational videos about their work.
Start date: 1 March 2008
End date: 31 August 2012
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website
lise-logo.png
LISE - Legal Language Interoperability Services
There is an urgent need for consolidated administrative nomenclatures and legal terminologies as tools to enhance interoperability and cross-border collaboration. Without high quality and standards-based terminologies, it is impossible to reach precision, efficiency, and transparency within and across any services, processes and systems in the areas of legal and administrative work.
Start date: 1 February 2011
End date:
31 July 2013
Project officer: Susan Fraser
Factsheet on EUROPA
Factsheet
Website
ltcompass-logo.png
LTCompass - Guiding Language Technology Paths from Research to Market
COMPASS is a Support Action for the Language Technology European research community and industry, that aims to: 1.Promote faster, wider and smoother market take-up of Language Technology (LT) research results by facilitating closer collaboration between research and industry players, and by generating better knowledge and understanding of demand/user needs and expectations 2.Foster the process of consolidation of LT stakeholders as a self-sustainable and innovative global industry, strategically co-ordinated under a common vision and approach to consumers' needs, and acting upon a shared innovation agenda for the continuous improvement of their technological base.
Start date: 1 November 2011
End date:
28 February 2014
Project officer: Kimmo Rossi
Factsheet on CORDIS
Factsheet
Website
LTWeb - Language Technology in the Web
The goal of LT-Web is to set the foundation for the integration of language technologies into core Web technologies, via the creation of a standard defining three kinds of metadata about: 1) information in Web content being relevant for language technology processing; 2) processes for creating Web content via localisation and content management work flows; 3) language technology applications and resources used in these applications.
Start date: 1 January 2012
End date:
31 December 2013
Project officer: Kimmo Rossi
Factsheet on CORDIS
Factsheet
Website

M

Back to top

mantra.png
MANTRA - Multilingual Annotation of Named Entities and Terminology Resources
This project will provide multilingual terminologies and semantically annotated multilingual documents, e.g., patent texts, to improve the accessibility of scientific information from multilingual documents.
Start date: 1 July 2012
End date: 30 June 2014

Project officer: Saila Rinne
Factsheet on CORDIS
Website
matecat-logo.jpg
MateCat - Machine Translation Enhanced Computer Assisted Translation
Worldwide demand of translation services has dramatically accelerated in the last decade, as an effect of the market globalization and the growth of the Information Society. Computer Assisted Translation (CAT) tools are currently the dominant technology in the translation and localization market. These include spell checkers, terminology managers, electronic dictionaries, full-text search tools, concordancers, bitexts, translation memory (TM) managers, and machine translation (MT) engines. Recent achievements by the so- called statistical MT approach have raised new expectations in the translation industry. So far, statistical MT has focused on providing ready-to-use translations, rather than outputs that minimize the effort of a human translator. The MateCat project aims at pushing what can be considered the new frontier of CAT technology: how to effectively integrate statistical MT within the translation workflow.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Factsheet
Website
MEDAR - Mediterranean Arabic Language and Speech Technology
MEDAR addresses International Cooperation between the EU and the Mediterranean region on Speech and Language Technologies for Arabic. The development of language resources and tools for the Arabic language will promote exchanges in the fields of culture and economy. By focussing on Arabic language technology and making both the technology and content available in Arabic helps to address the needs of citizens and make Arabic more accessible to the non-Arabic world, in order to stimulate business, research and mutual understanding.
Start date: 1 February 2008
End date: 31 July 2010
Project officer: Kimmo Rossi
Website

METALOGUE - Multiperspective Multimodal Dialogue: dialogue system with metacognitive abilities
The goal of METALOGUE is to produce a multimodal dialogue system that is able to implement an interactive behaviour that seems natural to users and is flexible enough to exploit the full potential of multimodal interaction. It will be achieved by understanding, controlling and manipulating system's own and users' cognitive processes. The new dialogue manager will incorporate a cognitive model based on metacognitive skills that will enable planning and deployment of appropriate dialogue strategies. The system will be able to monitor both its own and users' interactive performance, reason about the dialogue progress, guess the users' knowledge and intentions, and thereby adapt and regulate the dialogue behaviour over time.
Start date: 1 November 2013
End date: 31 October 2016
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website
meta-net-logo.jpg
META-NET (T4ME) - Technologies for the Multilingual European Information Society
Linguistic diversity is a corner stone of our multicultural European society. To preserve this essential asset in the age of the emerging information and knowledge society, Europe needs ICT technologies and applications at affordable costs that enable communication, collaboration and participation across language boundaries, secure their language users equal access to the information and knowledge society, and support each language in the advanced functionalities of networked ICT.
Start date: 1 February 2010
End date:
31 January 2013
Project officer: Kimmo Rossi
Factsheet on CORDIS
Factsheet
Website
metanet4u-logo.png
METANET4U - Enhancing the European Linguistic Infrastructure
METANET4U aims to contribute to the establishment of a pan-European digital platform that makes available language resources and services, encompassing both datasets and software tools, for speech and language processing, and supports a new generation of exchange facilities for them.
Start date: 1 February 2011
End date:
31 January 2013
Project officer: Kimmo Rossi
Factsheet on EUROPA
Factsheet
Website
metanord-logo.png
META-NORD - Baltic and Nordic Parts of the European Open Linguistic Infrastructure
META-NORD aims to establish an open linguistic infrastructure in the Baltic and Nordic countries by providing a description of the national landscape in terms of language use, contributing to a pan-European digital resource exchange facility, building and operating interconnected repositories and moblizing national and regional actors.
Start date: 1 February 2011
End date:
31 January 2013
Project officer: Kimmo Rossi
Factsheet on EUROPA
Factsheet
Website
mico logo
MICO - Media in Context
With the tremendous increase in multimedia content on the Web and in corporate intranets, discovering hidden meaning in raw multimedia is becoming one of the biggest challenges. Analysing multimedia content is still in its infancy, requires expert knowledge, and the few available products are associated with excessive price tags, while still not delivering sufficient quality for many tasks. This makes it almost impossible for normal companies, particularly SMEs, to make use of this technology. Also, analysis components typically operate in isolation and do not consider the context (e.g. embedding text) of a media resource.
Start date: 1 November 2013
End date: 31 October 2016
Project officer: Susan Fraser
Factsheet on CORDIS
Website

Mimics - multimodal immersive motion rehabilitation with interactive cognitive systems
The main hypothesis of this project is that movement training for neurorehabilitation can be substantially improved through immersive and multimodal sensory feedback. The approach is real-time acquisition of behavioural and physiological data from patients and the use of this to adaptively and dynamically change the displays of an immersive virtual reality system, with the goal of maximising patient motivation. This will result in complex systems that are natural, user-friendly and easy to use.
Start date: 1 January 2008
End date: 31 December 2010
Project officer: Philippe Gelin
Factsheet on CORDIS
Website

MLi - Towards a MultiLingual Data Services infrastructure
This Support Action will deliver the strategic vision and operational specifications needed for building the MLi (European MultiLingual data & services Infrastructure), formulate a multiannual plan for its development and deployment, and shape the multi-stakeholders alliances ensuring its long term sustainability.
Start date: 1 November 2013
End date:
31 October 2015
Project officer: Kimmo Rossi
Factsheet on CORDIS
Website
molto-logo.gif
MOLTO - Multilingual On-Line Translation
MOLTO's goal is to develop a set of tools for translating texts between multiple languages in real time with high quality. Languages are separate modules in the tool and can be varied; prototypes covering a majority of the EU's 23 official languages will be built.
Start date: 1 March 2010
End date:
28 February 2013
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website
monnet-logo.jpg
MONNET - Multilingual Ontologies for Networked Knowledge
The Monnet project will provide a semantics-based solution for integrated information access across language barriers, which is of growing importance to industry as exemplified by the Monnet use cases. A key solution to this problem is to deal with information at the semantic level, i.e. by abstracting away over language and form, allowing for more advanced and uniform: i) integration, ii) aggregation, iii) querying and iv) presentation of information across languages.
Start date: 1 March 2010
End date:
28 February 2013

Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
mormed-logo.jpg
MORMED - Multilingual Organic Information Management in the Medical Domain
MORMED proposes a multilingual community platform combining Web 2.0 social software applications with semantic interpretation of domain relevant content, enhanced with automatic translation capabilities, fine-tuned for a specific domain. MORMED will be piloted upon the community interested in Lupus or Antiphospholipid Syndrome (Hughes Syndrome), involving researchers, medical doctors, general practitioners, patients and patient support groups.
Start date: 1 April 2010
End date:
31 August 2012
Project officer: Susan Fraser
Factsheet on Europa
Factsheet
Website
mosescore-logo.png
MosesCore - Moses Open Source Evaluation and Support Co-Ordination for Outreach and Exploitation
Machine translation (MT) has increasingly become an indispensable tool for coping with Europes linguistic diversity. However machine translation systems are complex software applications requiring significant levels of cooperation and coordination to enable research to continue to flourish, and to improve the usage rate amongst commercial translators, and public bodies.
Start date: 1 February 2012
End date:
31 January 2015
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Factsheet
Website
multilingualweb-logo.png
MultilingualWeb - Advancing the Multilingual Web, Thematic Network
Given the importance of the World Wide Web to communication in all walks of life, and as the share of English Web pages decreases and that of languages spoken in the European Union increases, the importance of ensuring the multilingual viability of the World Wide Web is paramount. In order to build on current internationalization of the Web and move it forward, it is important to raise awareness of existing best practices and standards related to managing content on the multilingual Web, and look forward to what remains to be done. The project is coordinated by the World Wide Web Consortium (W3C), an organization of currently over 400 members worldwide from research and industry, headed by the Web's inventor, Sir Tim Berners-Lee. The other partners represent stakeholders from a range of affected areas.
Start date: 1 April 2010
End date:
31 March 2012
Project officer: Kimmo Rossi
Factsheet on EUROPA
Factsheet
Website
multisensor logo
MULTISENSOR - Mining and Understanding of multilinguaL contenT for Intelligent Sentiment Enriched coNtext and Social Oriented inteRpretation
The consumption of large amounts of multilingual and multimedia content regardless of its reliability and cross-validation can have important consequences on the society. An indicative example is the current crisis of the financial markets in Europe, which has created an extremely unstable ground for economic transactions and caused insecurity in the population. The fact that the national mass media provide exaggerated and contradictory information and the inaptness to understand local contexts from different countries has a considerable share in the aggravation of the crisis. To break this spiral, we need multilingual technologies with sentiment, social and spatiotemporal competence that are able to interpret, relate and summarize economic information and news created from various local subjective and biased views and disseminated via TV, radio, mass media websites and social media.
Start date: 1 November 2013
End date: 31 October 2016
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Website

O

Back to top

openerlogo.png
OpeNER - Open Polarity Enhanced Named Entity Recognition
Currently there are a multitude of companies offering Content Analytics and Social Internet Mining services for the purposes of Opinion Mining and Sentiment Analysis. Truly effective Sentiment Analysis is a complex NLP task in monolingual contexts alone. In multilingual contexts the complexity increase many-fold and also presents the challenge of comparison of opinion across languages and cultures. Named Entity Recognition and Classification is also key element to this challenge. As with is often the case in many innovative sectors and industries, a high percentage of SMEs are active offering niche solutions to specific segments of the market and/or domains. Acquiring or developing the base qualifying technologies require to enter this market is and expensive undertaking that redirects limited the resources of SMEs away from offering products and services that the market demands. The OpeNER project has as the goal of reuse and repurposing of exiting language resources and data sets to provide a set of underlying technologies to the broader community. OpeNER will focus on the provision of a supplementary sentiment lexicon with culturally normalised and graduated values.
Start date: 1 July 2012
End date: 30 June 2014
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website
organic-logo.png
ORGANIC - Self-organised recurrent neural learning for language processing
The human brain is an unrivalled “engine” for speech processing and language understanding. It integrates a large variety of learning, adaptation, optimization and self-stabilization mechanisms across many dynamically interacting levels of processing. The result of this highly entwined mesh of processes is supreme robustness, efficiency, and versatility. ORGANIC adopts principles of cortical architecture and self-organizing neurodynamics for the design of a new type of cognitive architectures for linguistic processing tasks. The consortium brings together pioneers in recurrent neural network research, cortical architectures for speech and language processing, speech processing research and an industrial partner who is leading in text recognition.
Start date: 1 April 2009
End date:
31 March 2012
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
organiclingua-logo.jpg
Organic.Lingua - Demonstrating the potential of a multilingual web portal for Sustainable Agricultural & Environmental Education
The Organic.Lingua project will extend the current Organic.Edunet portal to fill gaps in multilingual support and cross-language resource organisation and search, by significantly expanding its linguistic coverage. By doing so it aims to cover both the public sector needs by providing agricultural and environmental researchers and educators with a pan-European information, communication and collaboration platform, as well as by creating new business opportunities for the relevant private sector by demonstrating how a commercial system that will serve a global education market can be deployed.
Start date: 1 March 2011
End date:
28 February 2014
Project officer: Susan Fraser
Factsheet on EUROPA
Factsheet

P

Back to top

panacea
PANACEA - Platform for Automatic, Normalized Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies
A strategic challenge for Europe in today's globalised economy is to overcome language barriers through technological means. In particular, Machine Translation systems are expected to have a significant impact in the managing of multilingualism in Europe. PANACEA is facing the most critical aspect for Machine Translation to produce this expected impact in Europe: the so-called resources bottleneck. MT technologies are in themselves language independent, but they are inherently tied to the availability of language dependent resources.
Start date: 1 January 2010
End date:
31 December 2012
Project officer: Susan Fraser
Factsheet on CORDIS
Factsheet
parlance-logo.jpg
PARLANCE - Probabalistic Adaptive Real-Time Learning and Natural Conversational Engine
The project goal is to design and build mobile applications that approach human performance in conversational interaction, specifically in terms of the interactional skills needed to do so. These skills will include recognising and generating conversational speech incrementally in real-time, adapting to new concepts without manual intervention, and personalising interaction. All of these skills will be learned or adapted using real data, and will be used to build systems for interactive hyper-local search in three languages (English, Spanish and Mandarin) and for two domains such as property search and tourist information.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website
pheme logo
PHEME - Computing Veracity Across Media, Languages, and Social Networks
Social media poses three major computational challenges, dubbed by Gartner the 3Vs of big data: volume, velocity, and variety. Content analytics methods have faced additional difficulties, arising from the short, noisy, and strongly contextualised nature of social media. In order to address the 3Vs of social media, new language technologies have emerged, e.g. using locality sensitive hashing to detect breaking news stories from media streams (volume), predicting stock market movements from microblog sentiment (velocity), and recommending blogs and news articles based on user content (variety). PHEME will focus on a fourth crucial, but hitherto largely unstudied, challenge: veracity. It will model, identify, and verify phemes (internet memes with added truthfulness or deception), as they spread across media, languages, and social networks.
Start date: 1 January 2014
End date: 31 December 2016
Project officer: Susan Fraser
Factsheet on CORDIS
Website

PinView - Personal Information Navigator Adapting Through Viewing
The PinView consortium combines pioneering application expertise with a solid machine learning background in content-based information retrieval. the project will focus on developing new information retrieval principles needed for replacing or complementing explicit search queries. The research will facilitate a prototype of a proactive personal information navigator that allows retrieval of multimodal information (still images, text, video) available on the web and versatile databases. During browsing and searching with a task-dependent interface, the goals of the user will be inferred from explicit and implicit feedback signals and interaction (eye movements, pointer traces and clicks, speech) complemented with social filtering. The collected rich multimodal responses from the user are processed with new advanced machine learning methods to infer the implicit topic of the user's interest as well as the sense in which it is interesting in the current context.
Start date: 1 January 2008
End date: 31 March 2011
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website
pluto-logo.jpg
PLuTO - Patent Language Translations Online
PLuTO will build on existing state-of-the-art machine translation tools, currently successfully used for trademarks, and adapt them to the wider area of IP protection, and to patent translations in particular. Based on the experience of the existing consortium members, an online machine translation system will be built that is capable of assisting patent searchers with their multi-lingual information needs, much more reliably than general-purpose MT tools and much faster than human-based translations.
Start date: 1 April 2010
End date:
31 March 2013
Project officer: Susan Fraser
Factsheet on EUROPA
Factsheet
Website
portdiallogo.png
PortDial- Language Resources for Portable Multilingual Spoken Dialogue Systems
The PortDial project brings together European SMEs that are developing state-of-the-art spoken dialogue systems (SDS) and the handcrafted semantic components with research institutions at the forefront of progress in the automatic creation or enrichment of semantic language resources. The aim is to apply these technologies towards the creation of domain-specific multilingual SDS resources, specifically, data-linked ontologies and grammars.
Start date: 1 June 2012
End date: 31 May 2014
Project officer: Stefano Bertolo
Factsheet on CORDIS
Factsheet
Website
presemt-logo.png
PRESEMT - Pattern REcognition-based Statistically Enhanced MT
PRESEMT is a flexible and adaptable MT system, based on a language-independent method, whose principles ensure easy portability to new language pairs. This method attempts to overcome well-known problems of other MT approaches, e.g. bilingual corpora compilation or creation of new rules per language pair. PRESEMT will address the issue of effectively managing multilingual content and is expected to suggest a language-independent machine-learning-based methodology.
Start date: 1 January 2010
End date:
31 December 2012
Project officer: Michel Brochard
Factsheet on CORDIS
Factsheet
Website
prometheus.jpg
Prometheus - prediction and interpretation of human behaviour based on probabilistic structures and heterogeneous sensors
The project intends to establish a link between fundamental sensing tasks and automated cognition processes that concern the understanding a short-term prediction of human behaviour as well as complex human interaction. The analysis of human behaviour is unrestricted environments, including localization and tracking of multiple people and recognition of their activities, currently constitutes a topic of intensive research in the signal processing and computer vision communities.
Start date: 1 January 2008
End date: 31 December 2010

Project officer: Michel Brochard
Factsheet on CORDIS
Website
promislingua-logo.png
PROMISLingua - Performance Operational and Multilingual Interactive Services to support Compliance for SMEs in Europe
PROMISLingua's objectives are the translation, localisation and rollout of the existing PROMIS® online service with a total of eight languages in order to deliver a cost-efficient and easy-to-use internet-based service enabling SMEs to comply with safety, health, environment and quality regulations.
Start date: 1 April 2011
End date:
30 September 2013
Project officer: Pierre-Paul Sondag
Factsheet on EUROPA
Factsheet
Website

Q

Back to top

qtlplogo.png
QTLaunchPad - Preparation and Launch of a large-scale action for Quality Translation Technology
The support action will prepare the grounds for a new type of collaborative MT research dedicated to overcoming existing quality barriers. QTLaunchPad will assemble and provide needed data and tools in-cluding specialised translation corpora, test suites and tools for quality assessment, create a shared quality metrics for human and machine translation, improve automatic translation quality estimation, extend an existing platform for resource-sharing to the needs of quality MT research, define strategies and challenges and then plan and launch a large-scale research and innovation action for a breakthrough in quality translation technology. European research has the potential to overtake Google in the pursuit of automatic quality translation!
Start date: 1 July 2012
End date: 30 June 2014
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Factsheet
Website

QTLeap - Quality Translation by Deep Language Engineering Approaches
The support action will prepare the grounds for a new type of collaborative MT research dedicated to overcoming existing quality barriers. QTLaunchPad will assemble and provide needed data and tools in-cluding specialised translation corpora, test suites and tools for quality assessment, create a shared quality metrics for human and machine translation, improve automatic translation quality estimation, extend an existing platform for resource-sharing to the needs of quality MT research, define strategies and challenges and then plan and launch a large-scale research and innovation action for a breakthrough in quality translation technology. European research has the potential to overtake Google in the pursuit of automatic quality translation!
Start date: 1 November 2013
End date: 31 October 2016
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Website

R

Back to top

robocast_logo.gif
Robocast - robot and sensors integration as guidance for enhanced computer assisted surgery and therapy
The ROBOCAST project aims to develop ICT scientific methods and technologies which focus on robot assisted keyhole neurosurgery. A modular system, allowing a reduction of the footprint, will be developed with two robots and one active bio-mimetic probe, able to cooperate among themselves in a biomimetic sensory-motor integrated framework. A gross positioning 3-axes robot will support a miniature parallel robot holding the probe to be introduced through a "keyhole" opening into the skull of the patient. Optical trackers (tracking the end effector and the patient), an imaging endoscope camera, and electromagnetic position and force sensors (on the probe) will extend robot perception by providing the control system with position and force feedback from the operating tools, and with visual information of the surgical field.
Start date: 1 January 2008
End date: 31 December 2010
Project officer: Michel Brochard
Factsheet on CORDIS
Website

ROCKIT - Roadmap for Conversational Interaction Technologies
ROCKIT is a strategic roadmapping proposal for research and innovation in the area of natural conversational interaction. The primary scientific focus concerns interactive agents which are proactive, multimodal, social, and autonomous. A second focus concerns systems which can extract and exploit rich context and knowledge from heterogenous data sources. The main goal of ROCKIT is the development of a Research and Innovation Roadmap which integrates the vision and innovation agendas of those organisations (concerned with R&D and exploitation) in the field across Europe, with a broad coverage across sectors. A key goal is to bring together public sector research organisations with commercial organisations at all scales, with a particular focus on SMEs that represent the majority of fragmented commercial activity in Europe.
Start date: 1 December 2013
End date: 30 November 2015
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website

S

Back to top

savas-logo.png
SAVAS - Sharing AudioVisual Language Resources for Automatic Subtitling
Nowadays, due to the quantity of the demand and the cost of the process, manual subtitling is no longer feasible. Broadcasters and subtitling companies are seeking for more productive subtitling alternatives. In this context, SAVAS partners aim to acquire, share and reuse audiovisual resources of broadcasters and subtitling companies so that high-tech European ASR (Automatic Speech Recognition) companies can use the shared data to develop domain specific Large Vocabulary Continuous Speech Recognisers (LVCSRs) in new languages to solve the automated subtitling needs of the media industry. Within the project, data and LVCSR technology for automated subtitling will be collected, shared and developed for the following six languages: Basque, Spanish, Italian, French, German and Portuguese.
Start date: 1 May 2012
End date: 30 April 2014
Project officer: Susan Fraser
Factsheet on CORDIS
Factsheet
Website
semaine-logo-small.jpg
SEMAINE - sustained Emotionally colooured Machine-human Interaction using Nonverbal Expression
The aim of the SEMAINE project is to build a Sensitive Artificial Listener – a multimodal dialogue system with the social interaction skills needed for a sustained conversation with a human user. The system will emphasise “soft” communication skills, i.e. non-verbal, social and emotional perception, interaction and behaviour capabilities. The Sensitive Artificial Listener paradigm involves only very limited verbal capabilities, but has been shown to be suited for prolonged human-machine interaction. In this paradigm, we will build a real-time, robust interactive system perceiving a human user's facial expression, gaze, and voice, and engaging with the user through an Embodied Conversational Agent's body, face and voice. The agent will exhibit audiovisual listener feedback in real time while the user is speaking, and will take the user's feedback into account while the agent is speaking. The agent will pursue different dialogue strategies depending on the user's state; it will learn to interpret the user's non-verbal behaviour and adapt its own behaviour accordingly.
Start date: 1 January 2008
End date: 31 December 2010
Project officer: Philippe Gelin
Factsheet on CORDIS
Website

SENSEI - Making Sense of Human-Human Conversation Data
The overall goals of the SENSEI project are twofold. First, SENSEI will develop summarization/analytics technology to help users make sense of human conversation streams from diverse media channels. Second, SENSEI will design and evaluate its summarization technology in ecological environments, aiming to improve task performance and productivity of end-users. Conversational interaction is the most natural and persistent paradigm for business relations with end-customers or users. In contact centres millions of customer spoken conversations are handled daily. On social media platforms hundreds of millions of blog posts are delivered through generalist or proprietary platforms. In both cases, conversations have little impact on the intended target "listeners" due to the volume, velocity and diversity (media, style, social context) of the document streams (spoken conversations and blog posts). Most language analytics technology is limited in that it performs keyword search, which does not provide automatic descriptions of what happened, who said what, which opinions are held on what subject, in a coherent, readable and executable form.
Start date: 1 November 2013
End date: 31 October 2016
Project officer: Carola Carstens
Factsheet on CORDIS
Website
sera-logo.jpg
SERA - Social Engagement with Robots and Agents
The project SERA aims to advance science in the field of social acceptability of verbally interactive robots and agents, with a view especially to their applications in assistive technologies (companions, virtual butlers). To this aim, the project will undertake a field study in three iterations to collect data of real-life, long-term and open-ended relationships of subjects with robotic devices. The three iterations test different conditions (functionalities) of the equipment, which will consist of a room equipped with sensors at the subjects' home, a computer and a simple robotic device (the Nabaztag) as the front-end for interaction. The project partners will analyse the collected audio and video data in parallel, using different, mainly qualitative, methods. Data analysis will be prepared and accompanied by theoretical and methodological research in order to a) take into account the state of the art and b) ensure the quality of the field study. The project will use findings from the field study to specify, build and implement a reference architecture for social engagement, and use it for developing a showcase system of combined speech based service applications with relevance to the target field and audience.
Start date: 1 January 2009
End date:
31 December 2010
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Website
signspeak-logo-new.jpg
SignSpeak - Scientific understanding and vision-based technological development for continuous sign language recognition and translation
Deaf communities revolve around sign languages as they are their natural means of communication. Although deaf, hard of hearing and hearing signers can communicate without boundaries amongst themselves, there is a serious challenge for the deaf community in trying to integrate into educational, social and work environments, as the vast majority of Europeans do not have signing skills. The overall goal of SignSpeaker is to develop a new vision-based technology for translating continuous sign language to text, in order to improve the communication between deaf and hearing communities. This is the second FP7 project addressing sign language.
Start date: 1 April 2009
End date:
31 March 2012
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Factsheet
Website
simple4all-logo.jpg
SIMPLE4ALL - Speech synthesis that improves through adaptive learning
The Simple4All project will create speech synthesis technology that learns from data with little or no expert supervision and continually improves itself, simply by being used. In order to be accepted by users, the voice of a spoken interaction system must be natural and appropriate for the content. Using the same voice for every application is not acceptable to users. But creating a speech synthesiser for a new language or domain is too expensive, because current technology relies on labelled data and human expertise.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Leonhard Maqua
Factsheet on CORDIS
Factsheet
Website
sspnetlogo.gif
SSPNET - Social Signal Processing Network
The ability to understand and manage social signals of a person we are communicating with is the core of social intelligence. Social intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. Although each one of us understands the importance of social signals in everyday life situations, and in spite of recent advances in machine analysis and synthesis of relevant behavioural cues like blinks, smiles, crossed arms, laughter, etc., the research efforts in machine analysis and synthesis of human social signals like empathy, politeness, and (dis)agreement, are few and tentative.
Start date: 1 February 2009
End date: 31 January 2014
Project officer: Philippe Gelin
Factsheet on CORDIS
Website
sumat.gif
SUMAT - an online service for subtitling by machine translation
Subtitling plays a very important role as it is the preferred multimedia content translation method in most of the European countries and for most of the genres in order to make audiovisual content widely accessible across languages. The increasing use and transmission of digital multimedia and multilingual content through the Web, DVDs and the current European and national policis promoting the subtitling of content broadcast by public TV stations have had as a consequence that subtitling demands have increased in recent years.
Start date: 1 April 2011
End date:
31 March 2013
Project officer: Susan Fraser
Factsheet on EUROPA
Factsheet
Website

T

Back to top

taaslogo.png
TaaS - Terminology as a Service
The aim of the TaaS project is to create a cloud-based platform for acquiring, cleaning up, sharing, and reusing multilingual terminological data, one of the most important language resources for industry, academia, and society in general. The motivation of the TaaS project is to address an evident need for instant access to the most recent terms and direct user involvement in the creation and sharing of terminology data. Such a platform will provide a variety of online services for key terminology tasks becoming an integral part of the multifaceted global cloud-based service infrastructure. TaaS services will perform term identification in the user provided documents. The platform will perform terminology extraction and return a list of term candidates.
Start date: 01 June 2012
End date: 31 May 2014
Project officer: Aleksandra Wesolowska
Factsheet on CORDIS
Website
translectures-logo.png
transLectures - Transcription and Translation of Video Lectures
Online educational repositories of video lectures are rapidly growing on the basis of increasingly available and standardised infrastructure. A well-known example of this is the VideoLectures web portal, a free and open access educational video lectures repository, and a major player in the development of the widely used Opencast Matterhorn platform for educational video management. As in other repositories, transcription and translation of video lectures in VideoLectures is needed to make them accessible to speakers of different languages and to people with disabilities. However, also as in other repositories, most lectures in VideoLectures are neither transcribed nor translated because of the lack of efficient solutions to obtain them at a reasonable level of accuracy. The aim of transLectures is to develop innovative, cost-effective solutions to produce accurate transcriptions and translations in VideoLectures, with generality across other Matterhorn-related repositories.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Susan Fraser
Factsheet on CORDIS
Factsheet
Website
trendminer-logo.png
TrendMiner - Large Scale, Cross-lingual Trend Mining and Summarisation of Real-Time Media Streams
The recent massive growth in online media and the rise of user-authored content (e.g weblogs, Twitter, Facebook) has lead to challenges of how to access and interpret these strongly multilingual data, in a timely, efficient, and affordable manner. Scientifically, streaming online media pose new challenges, due to their shorter, noisier, and more colloquial nature. Moreover, they form a temporal stream strongly grounded in events and context. Consequently, existing language technologies fall short on accuracy, scalability and portability.
Start date: 1 November 2011
End date:
31 October 2014
Project officer: Susan Fraser
Factsheet on CORDIS
Factsheet
Website
ttc-logo-small.jpg
TTC - Terminology Extraction, Translation Tools and Comparable Corpora
The TTC project aims at improving machine translation tools, computer-assisted translation tools and multilingual content management tools by automatically generating bilingual terminologies from comparable corpora in several European languages (i.e. English, French, German and Latvian), as well as in Chinese and Russian.
Start date: 1 January 2010
End date:
31 December 2012
Project officer: Pierre-Paul Sondag
Factsheet on CORDIS
Factsheet
Website

X

Back to top

xlike-project-logo.png
X-LIKE - Cross-Lingual Knowledge Extraction
The goal of the X-LIKE project is to develop technology to monitor and aggregate knowledge that is currently spread across global mainstream and social media, and to enable cross-lingual services for publishers, media monitoring and business intelligence. In terms of research contributions, the aim is to combine scientific insights from several scientific areas to contribute in the area of cross-lingual text understanding. By combining modern computational linguistics, machine learning, text mining and semantic technologies we plan to deal with the following two key open research problems: - to extract and integrate formal knowledge from multilingual texts with cross-lingual knowledge bases, and - to adapt linguistic techniques and crowdsourcing to deal with irregularities in informal language used primarily in social media.
Start date: 1 January 2012
End date:
31 December 2014
Project officer: Susan Fraser
Factsheet on CORDIS
Factsheet
Website

xLiMe – crossLingual crossMedia knowledge extraction
Europe is different from other large media markets such as the US or China in that information is being generated in different languages and distributed via diverse streams of localised media channels. Automatic analysis is complicated further by different content types (audio, video, text) and different channels (mainstream, social media). Thus, information can only be analysed independently for each dimension. This restricts the extractable knowledge and keeps it fragmented, which ultimately constrains the exchange of information. xLiMe proposes to extract knowledge from different media channels and languages and relate it to cross-lingual, cross-media knowledge bases.
Start date: 1 November 2013
End date:
31 October 2016
Project officer: Saila Rinne
Factsheet on CORDIS
Website
Source: http://cordis.europa.eu

Change segments status in SDL Trados

Select all the segments where you want to change the status.
To do this, place the cursor in the first segment that you want to change back to untranslated. Make sure it's in the segment number column on the left.  Now go the end of the file (ctrl+end) and click shift+left mouse click in the last segment number.  With all the segments selected, do a right mouse click and select "change segment status" to "not translated"

The Best Free Dictionary and Thesaurus Programs and Websites


00_lead_image_dictionaries
Are you a writer? Or a word geek? If you write anything or play word games, dictionaries, thesauruses, and other reference tools can come in handy. We’ve found some useful offline and online tools for looking up words, finding synonym, or building words in Scrabble.
Dictionary and thesaurus programs and websites allow us to go beyond the dated, printed dictionary. You can find so much more up-to-date information on the web without having to buy dictionaries and other reference books. The following programs and websites provide useful reference tools for free.

Offline Desktop Software

The following software programs are offline programs you download and install on your computer. Some of the downloads may be large because the programs include the full dictionaries with the program for offline use.

Ultimate Dictionary

Ultimate Dictionary is a free dictionary program for Windows that is easy to install and use. It contains a collection of around 61 dictionaries, which includes dictionaries, thesauruses, and glossaries, as well as English, Spanish, French, and Polish word references.
When you type a word into the search box Ultimate Dictionary, it looks for it in all the dictionaries it contains at once and provides a meaning from each dictionary in which the word was found in the right pane. The results start to display as you type the word with hints provided for matching words. Quickly jump from one information source to another (translations, encyclopedia entries, glossaries) simply by clicking the Jump to Dictionary button and selecting a reference source from the list on the left.
ultimate_dictionary

WordWeb

WordWeb is a free dictionary program for Windows that provides a dictionary and a thesaurus. Each word is accompanied by synonyms, and where available antonyms (a word having a meaning opposite to that of another word), meronyms (a word that names a part of a larger whole; for example, “finger” is a meronym of “hand”), hypernyms (a word that is more generic than a given word; for example, “car” or “airplane” are hypernyms for more precise terms like Toyota Camry or Boeing 747), holonyms (a word that names the whole of which a given word is a part; for example, “hat” is a holonym for “brim” and “crown”).
WordWeb provides a keyboard shortcut that allows you to look up any word anywhere on your computer. By default, the keyboard shortcut is Ctrl + Right Click, but this can be changed in the Preferences. WordWeb also has access to online references that include Wikipedia, Wiktionary, and WordWeb Online, allowing you to access word definitions on the internet.
You can also access a history or your word searches, copy the results of a lookup, and click on any word displayed in WordWeb to display a definition of that word.
WordWeb is free, but they have a very unique licensing policy that you must satisfy to continue using the free program. Otherwise, you must purchase WordWeb Pro, which includes 5000 more definitions and many extra features.
wordweb_free_orig

TheSage English Dictionary and Thesaurus

TheSage is a free offline dictionary and thesaurus that contains a large dictionary of more than 210,000 definitions, as well as a complete thesaurus providing over 1,400,000 relationships between definitions, including synonyms, antonyms, hyponyms, hypernyms, meronyms, holonyms, and more.
TheSage has a tabbed interface, allowing you to look up multiple words at the same time. You can also search using wildcards and refer to a history of your searched words. The program also integrates into the Windows context menus, providing quick access to definitions from almost anywhere on your computer.
Other features include access to online references such as Wikipedia and Google (for definitions and translation), an anagram search, real-time search, system tray integration, and the ability to run TheSage from a USB flash drive.
the_sage

LingoPad

LingoPad is a free offline dictionary for Windows that contains a German-English dictionary, as well as other dictionaries that are available for download. You can import your own dictionaries, define your own, and create and edit word lists.
The search feature allows you to search for the beginning, ending, or middle part of a word, search for collocations, and view a list of the latest search words. You can also use a hotkey to search for a tagged word or a word from the clipboard. There are direct links to look up words on Wikipedia and various search engines.
NOTE: Currently, LingoPad is only available for Windows, but there will be no further development of LingoPad only for Windows in the near future. However, they are planning to develop a cross-platform version of LingoPad for Windows, Linux, and Mac OS X.
lingopad_orig

Artha Dictionary

Artha is a free, cross-platform (Windows and Linux), offline English Thesaurus based on WordNet, which is a large, lexical database of English developed at Princeton University. For each word you look up, the possible relatives displayed by Artha include Synonyms, Antonyms, Derivatives, Pertainyms (related noun/verb), Attributes, Similar Terms, Domain Terms, Entails (what a verb entails doing), Causes (what a verb causes to), Hypernyms (is a kind of), Hyponyms (kinds), Holonyms (is a part of) and Meronyms (parts). Each of these links takes you to a definition and example of the relative category. When you select a relative in Artha, its corresponding definition is scrolled to and highlighted for easy understanding.
Artha starts searching as you start typing, displaying a list of possible matches from what you have typed so far. If you misspell a search word, Artha provides near-match suggestions.
Once you launch Artha, it sits in the system tray. You can look up words from any open window by selecting the text and pressing a pre-set hotkey. When using the hotkey, you can choose to have Artha show a passive notification, or balloon tip, so you can quickly continue what you were doing.
Artha maintains a history of all your search terms and provides Previous and Next buttons so you can easily browse through your searches.
artha_thesaurus_orig

Aard Dictionary

Aard Dictionary is a free offline dictionary and thesaurus program that allows you to look up meanings for difficult words using multiple dictionaries in multiple languages. You can add more than 50 different dictionaries to its database. Some of its features include on-the-fly searching and filtering capabilities, the ability to view a history or recent searches, the ability to zoom in and out of contents, and fast word lookups.
For words found on Wikipedia using Aard Dictionary, the full article is displayed and you can save the Wikipedia article as an HTML document on your local hard disk. Aard Dictionary also provides many keyboard shortcuts for those of you who prefer using the keyboard over the mouse for navigation.
aard_dictionary_orig

GoldenDict

GoldenDict is a free, feature-rich offline dictionary program for Windows and Linux. It supports multiple dictionary formats, such as Babylon, StarDict, and WordNet. You can also look up words on Wikipedia, Wiktionary, or any other MediaWiki-based websites using GoldenDict. It also supports looking up and listening to pronunciations using the pronunciation guide, forvo.com.
You can lookup words selected in other applications or translate a word from the clipboard using hotkeys.
goldendict_orig

Lingoes

Lingoes is a free, easy-to-use dictionary and text translation program that offers lookup dictionaries, full text translation and pronunciation of words in over 80 languages. You can also select text in other programs and windows and use a key combination to have Lingoes to display results for the selected text.
Lingoes provides free dictionaries, thesauruses, and encyclopedias for download and allows you to configure them yourself. In addition to using offline dictionaries, you can also use online dictionary services and encyclopedias such as Wikipedia. Lingoes also includes tools and references, such as a currency converter, international dialing codes, international time zone converter, weights and measures, and more. You can download additional tools, information, games, and shortcuts.
lingoes_orig

Mobysaurus Thesaurus

Mobysaurus Thesaurus is a free, offline, English thesaurus program that contains 30,260 roots and more than 2.5 million synonyms. It suggests results as you type, stores your look up history, and allows you to bookmark your favorite words and phrases. You can also perform a wildcard word search, search online for results, print your results, and copy your results to the clipboard.
Mobysaurus Thesaurus is freeware, but they ask you to register for a license key. It doesn’t cost anything. Register on their forum to receive a license key that is good for six months. After that, you need to go back to their site to get a new license key for another six months. One year after your initial signup, you can go back to the site and get a permanent, non-expiring license key.
Mobysaurus Thesaurus requires at least .NET Framework 2.0 to run.
NOTE: You can also download Mobysaurus Thesaurus from http://www.softpedia.com/get/Others/Home-Education/Mobysaurus.shtml, but you still need to register on DonationCoder.com’s site for a license key.
mobysaurus

CleverKeys

CleverKeys is a free tool that allows you to access definitions at Dictionary.com, synonyms at Thesaurus.com, facts at Reference.com, and more from just about any Windows or Mac OS X program, such word processors, web browsers, and email programs.
Some features of CleverKeys include the ability to add your own URLs, highly configurable hotkeys, and the ability to open CleverKeys results in a new window. CleverKeys works with most current web browsers, including Firefox, Chrome, and Internet Explorer.
cleverkeys

Online

The following websites provide online dictionaries and other reference material.

Dictionary.com

Dictionary.com is a free, online dictionary that is part of the Reference.com website (with Thesaurus.com, as well) and is an interactive and reliable source for word meaning and usage. With 50 million users around the world accessing the site every month, It’s the largest and most authoritative free online dictionary and mobile reference resource.
Dictionary.com can also be used on mobile platforms, including Android, iPhone, iPad, BlackBerry, Windows Phone, Nook, and Kindle Fire.
dictionary_dot_com

TheFreeDictionary.com

The Free Dictionary is the world’s most comprehensive, free, online dictionary with dictionaries available in many languages and other types of dictionaries available, such as Medical, Legal, and Financial dictionaries. It’s also contains a Thesaurus, Acronyms and Abbreviations, Idioms, Encyclopedia, a Literature Reference Library, and a search engine.
You can also make The Free Dictionary your own home page and customize it by adding and removing, dragging and dropping, and “using or losing” existing content windows. You can also add your own bookmarks, weather information, horoscope, and RSS feeds from anywhere on the internet.
You can access definitions on The Free Dictionary from almost any application on your PC by installing the free Desktop Assistant for Windows. Then, simply highlight a word, press Alt and left click your mouse, and the definition displays.
You can also install add-ons in Firefox, Chrome, Internet Explorer, and Safari that make access to The Free Dictionary quick and easy from almost any other webpage.
the_free_dictionary

Wiktionary

Wiktionary is a free, multilingual dictionary that is the lexical companion to Wikipedia. It includes a thesaurus, a rhyme guide, phrase books, language statistics and extensive appendices, as well as extra, clarifying information in the definitions, such as etymologies, pronunciations, sample quotations, synonyms, antonyms and translations.
Wiktionary is a wiki, so you can edit it. However, before you contribute, be sure to read through some of the Help pages, as Wiktionary does things differently than other wikis and they have some strict criteria.
wiktionary

Merriam-Webster Online Dictionary

The Merriam-Webster Online Dictionary is a free dictionary based on the print version of Merriam-Webster’s Collegiate® Dictionary, Eleventh Edition. The online dictionary includes the main A-Z listing of the Collegiate Dictionary, as well as the Abbreviations, Foreign Words and Phrases, Biographical Names, and Geographical Names sections of that book.
merriam_webster

Oxford Dictionaries Online

Oxford Dictionaries Online is a free website that offers a comprehensive, current English dictionary, information to help you write better, puzzles and games, and a language blog. It also contains up-to date bilingual dictionaries in French, German, Italian, and Spanish.
They also offer a premium site, Oxford Dictionaries Pro (£42 per year), which features more than 350,000 definitions, 600,000 synonyms and antonyms, and over 1.9 million usage examples, all linked intelligently and seamlessly, audio pronunciations, example sentences, advanced search functionality, a puzzle solver section, and specialist language resources for writers and editors.
oxford_dictionaries

Collins Dictionary

The official online Collins Dictionary contains over 1 million entries, and provides free access to word definitions, translations, and examples. There is also a section where you can play word games, such as crossword puzzles and cryptic crossword puzzles, use their Scrabble Checker, and learn from their top 10 Crossword Tips and top 10 scrabble tips.
They also provide a widget that allows you to place the Collins English Dictionary, French Dictionary, Spanish Dictionary, German Dictionary or Italian Dictionary on your website or blog.
collins_american_english_dictionary

Word-Net Online

Word-Net Online is a free, English dictionary and thesaurus that offers definitions, synonyms, antonyms, hypernyms, hyponyms, meronyms, holonyms, derived forms and sample sentences.
wordnet_online
They also offer a free dictionary and thesaurus program for your Windows desktop. Install the program and search for any word or term on WordNet-Online directly from your desktop.
wordnet_online_dialog_orig

OneLook

OneLook is a free dictionary site that can be thought of as a search engine for words and phrases. When you look up a word on OneLook, you are directed to the web-based dictionaries that define or translate that word. The OneLook search engine indexes more than five million words in more than 1000 online dictionaries.
You can also use OneLook to find a word when you only have part of it, similar to a crossword dictionary. Enter a pattern using letters and the wildcards * and ? to display a list of words matching your pattern. If you enter a word and select Find translations, OneLook displays a list of dictionary websites that have translations of that word into other languages.
OneLook can be customized to the way you work, as well.
onelook
OneLook also has an interesting feature called a Reverse Dictionary. This allows you to describe a concept and receive a list of words and phrases related to that concept. You can enter a few words, a sentence, a question, or even one word. The best results are obtained if you keep the entered phrase short and the best matches are shown first in the list of results.
onelook_reverse_dictionary

Urban Dictionary

Urban Dictionary is a web-based dictionary of slang words and phrases written and rated by site visitors (voted thumbs “up” or “down”) and regulated by volunteer editors. As of September 6, 2012, it contained 6,743,306 definitions. Visitors to Urban Dictionary may submit definitions without registering, but they must provide a valid email address. All new definitions must be approved by the editors. Once approved, entries become the property of Urban Dictionary.
The definitions on Urban Dictionary are meant to be slang or ethnic culture words, phrases, and other terms not found in standard dictionaries. Most words have multiple definitions and include usage examples.
urban_dictionary

MetaGlossary.com

MetaGlossary is a comprehensive, highly organized dictionary and glossary. It’s the world’s largest, constantly-updated repository of information, harvesting definitions from the entire web. MetaGlossary precisely extracts the meanings of the terms and phrases it finds on the web and provides you with concise, direct explanations for the terms and phrases. The meanings are also organized based on topic and usage.
metaglossary

Specialty Dictionary Programs and Websites

In our search for useful offline and online dictionaries, thesauruses, and other reference tools, we also came across some dictionary programs and websites that serve special purposes, such as a dictionary only defining computer terms, a crossword dictionary, and a few Scrabble dictionaries and word finders.

Smadi Computer Dictionary

The Computer dictionary is reference tool with a very simple interface, in which you can find over 86,000 computer and technology related acronyms and their definitions.
computer_dictionary_orig

Crossinary

Crossinary is a dictionary program for Windows that helps you create and solve crosswords and other word puzzles. It contains over a quarter of a million English words (including some proper nouns) and allows you to quickly search the massive list to find all the words that match a pattern, such as g-something-something-k.
crossinary_orig

Free Scrabble Dictionary

The Free Scrabble Dictionary is a free tool to help you look up definitions and find words to use when playing Scrabble. They have tens of thousands of words and a half a million definitions. You can also use the Free Scrabble Dictionary to play other word games, such as Words with Friends, Scramble with Friends, and Hanging with Friends.
free_scrabble_dictionary
You can also use the Hasbro Scrabble Dictionary and Word Builder and the Word Finder websites to help you play word games, such as Scrabble®, Lexulous, Wordscraper, Scrabulous, Anagrammer, Literati, Text Twist, Jumble Words and Words with Friends.
scrabble_word_finder
If you’ve found any useful dictionary, thesaurus, or reference resources or other dictionaries and tools for playing word games, that are not on this list, let us know.
Source: howtogeek.com

Software with TBX support

This is a list of software claiming TBX support of some kind. It contains links to the software website, a brief description of the extent of its support of TBX, and other information (type, price, etc.). We are currently in the process of validating the TBX output of each of the listed products. The results of the validation will be shown in the final column as it is finished.
The widespread adoption of TBX would be a boon to the translation industry. By sharing a list of software with TBX support, we hope to further encourage its use both in further software development and personal data storage. Using TBX means easy data sharing between each of the products listed here: a non-trivial benefit!
Click the column headers to sort by that column.
Name and Website Tool Type Web-based? Nature of TBX Support Claims Price Compliance
Content Checker no Checks valid core structure, and then adherence to an XCS file. open source
Translation Environment Tool (TeNT) no
import/export
open source TBD
TeNT no import/export TBX-Basic only (see this document) $599 TBD
SDL Multiterm Terminology management no
import/export
$300
Translation Management System (TMS) no
import/export
open source TBD
Terminology management yes import/export ? TBD
TeNT no Convert from CSV or MARTIF to TBX
import/export
298-598 TBD
TeNT/TMS yes import/export
€0-199
TBD
Terminology Management optional import/export pro- $795
web- ?
web extension module: $4200
TBD
Terminotix Synchroterm Term extraction no export $420 TBD
Maxprograms Anchovy Terminology management/extraction no import/export "TBX with default XCS or proprietary extensions, cusomizable using XSL stylesheets" Swordfish: €260 (see also student/site prices) TBD
Mneme TeNT no import/export free TBD
Terminology Management (OWL) based no import (conversion to OWL) free TBD
Quality Assurance no import (only?)
€249-2500
TBD
Terminology management optional import/export suite-?
cloud-?
TBD
TMS yes
import/export/conversion (via Translate Toolkit) open source TBD
TEnT no
import/export/conversion (via Translate Toolkit) open source TBD
Toolkit/API for file conversion QA, and more general functions no
import/export/conversion open source TBD; see here for progress.
XBench QA and Terminology Management no import/export
€39/yr; varies by residency; beta is free
TBD
Lingotek TeNT yes import/export ? TBD
Metatexis CAT no import/export Word: $50-180
Server: ? (free for training purposes)
TBD
Okapi Checkmate QA no import/export open source TBD
TeNT/localization no import/export ? TBD
TeNT no import/export (via plugins)
€95-2900
TBD
Termbases Terminology Management yes import/export Personal-free
€30/mo 300/yr
TBD
Wordbee TeNT yes import (and export?) starts at $178/6 mo $323/yr TBD
Star Transit/TermStar CAT no ? ? TBD
Idiom Worldserver TeNT no import/export free? TBD
Across TeNT no import/export free for freelancers TBD
Wordfast TeNT no import/export
€200-500
TBD
Acrolinx Authoring/Terminology Management no import/export ? TBD
RC-WinTrans localization no imports Microsoft glossaries (which are TBX) $795-4575 TBD
Lingobit Localizer Enterprise localization no import/export $1950 TBD
localization/TeNT no import (and export?)
€620
TBD
TeNT cloud:yes import/export Cloud: free - $230/mo
Editor: free
Server: ?
TBD
JiveFusion CAT no import/export ? TBD
OpenTM2 TeNT no import/export open source TBD
Fluency TeNT/other optional import/export Translation suite starts at $349 TBD
Text United TeNT/other yes import/export (not for client's projects) single user- free

TBD
DVX2 TeNT no import/export 590 €

TBD
MultiTrans Prism TMS no import/export 590 €

TBD
Terminator Terminology Management yes import/export Open Source

TBD
QA-Distiller Quality Assurance no export Free using plugin

TBD
Weblate CAT yes import/export Free

TBD
The Microsoft Language Portal Downloads also deserve mention because they offer sizable terminology downloads in TBX format, for free.
Source: http://www.tbxconvert.gevterm.net

Removing frames (text boxes) from a word document, after OCR or saving as rtf from pdf document

You saved or scanned a document with OCR software like Abbyy FineReader or OmniPage Pro? You saved as rtf a PDF document and the resultant word document, contains multiple frames?

Frames make the document very hard to edit because all text is placed inside frames. We need to remove those frames if we want to edit the document.

How do we do that?

If you do not care about formatting you do this:

1.
—Open the file which has frames in MS Word
—Save the file as a Plain text file.
—Open the new text file you have just saved in Notepad or WordPad or some other text editor.
—Now Select all the text by pressing Ctrl+A, Copy and paste that into a New MS Word file. Then Save it with any name you want. Frames are gone.

If you do care about formatting:

2.
—Copy everything in the Word document, paste all the text into WordPad, copy all the text in the WordPad document, and paste it back into the Word document.

3.
—Select the entire document by pressing Ctrl+A, and then press Ctrl+Q. This will set every paragraph back to its default condition and most likely remove the frames.


4.

Use a macro to remove text boxes and delete text

Code: [Select]
Sub DeleteTextBoxesAndText()
Dim oShp As Word.Shape
Dim i As Long
For i = ActiveDocument.Shapes.Count To 1 Step -1
Set oShp = ActiveDocument.Shapes(i)
If oShp.Type = msoTextBox Then
oShp.Delete
End If
Next i
End Sub

5.

Use a macro to remove text boxes but keep text

Code: [Select]
Sub RemoveTextBox2()
    Dim shp As Shape
    Dim oRngAnchor As Range
    Dim sString As String

    For Each shp In ActiveDocument.Shapes
        If shp.Type = msoTextBox Then
            ' copy text to string, without last paragraph mark
            sString = Left(shp.TextFrame.TextRange.Text, _
              shp.TextFrame.TextRange.Characters.Count - 1)
            If Len(sString) > 0 Then
                ' set the range to insert the text
                Set oRngAnchor = shp.Anchor.Paragraphs(1).Range
                ' insert the textbox text before the range object
                oRngAnchor.InsertBefore _
                  "Textbox start << " & sString & " >> Textbox end"
            End If
            shp.delete
        End If
    Next shp
End Sub

6.

Use a macro to remove frames

Code: [Select]
Sub RemoveFrames()
    Dim aFrame As Frame
    Dim p As Paragraph
    Dim l As Single

    For Each aFrame In ActiveDocument.Frames
       aFrame.RelativeHorizontalPosition = wdRelativeHorizontalPositionPage
       l = aFrame.HorizontalPosition
       For Each p In aFrame.Range.Paragraphs
          p.LeftIndent = l
       Next p
       aFrame.Delete
    Next aFrame
End Sub

7.

Use a macro to remove text boxes but keep text (commercial tool, free trial)

Quickly remove all text boxes and keep texts in Word

and a macro for frames by the same tool:
Quickly remove all frames and keep text from document in Word

Source: https://www.translatum.gr

Goldendict for Linux Mint

Introduction

While spelling dictionaries of the likes of a-, i-, my- and hunspell were and are still common in the UNIX and Linux world, and in open source software in general, the usage of explanatory and bilingual dictionaries and also thesauri mostly requires proprietary software. This is probably because their editing is rooted in a tradition which is centuries old, whereas automatic spelling became only possible — and necessary — when text begun to be edited with computers.
Nevertheless, these kinds of reference dictionaries are indispensable if you need to translate some text, be it shorter or longer, in any given language outside a web browser, or when you’re offline. All in all, we have the following alternatives:
  1. online dictionaries which maybe reflect the modern usage of a spoken language, but which are neither complete nor accurate;

  2. free offline dictionaries which may be complete, but which are most likely to be free of copyright issues only when they’re really dated;

  3. non-free offline dictionaries which are both up to date and complete, but which ask for a paid license.
Hence, the best would be to find an optimal mix of all these resources, and a look-up program that could easily cope with all of them. The package named GoldenDict is such a dictionary viewer.

  The viewer

Actually, the viewer we’re going to use is a kind of a slim web browser which is based on the same WebKit rendering engine as Cinnamon. Its concept comes in handy when local dictionaries and web pages are to be displayed together as a result of a query. Once properly configured, we don’t need to bother where the sources are located, and even the choice of correct dictionaries will follow the session’s current language.
The GoldenDict package can be installed easily from the Software Manager. When we’ll be done with this tutorial, you’ll be able to look up any word or text fragment in all your relevant dictionaries at once like in the following picture, and this function will be at your fingertips with a single keypress from within any application if you take care to select that text beforehand:

GoldenDict screenshot

The very first step

At the very beginning, you should uncheck the Edit› Preferences› Scan Popup› Enable scan popup functionality, otherwise you’ll get tired of GoldenDict even before you start using it. That is to say, with the first dictionary just installed, you’ll be bugged by an unprompted popup as soon as you select some text. Thus better disable this — after all, we want the dictionary to help us and not to get into our way 




Sources tab  1. External tools

You’ll need three additional utilities which can all be installed from the Software Manager as usual:
  • xsel to manipulate the selected text;

  • espeak, a speech synthesizer to give an idea about the pronunciation of unknown words;

  • uzbl, a lightweight web browser to visualize contents like Google Translate which cannot be rendered inside a GoldenDict window.
It’s just the last one, the Uzable Browser that needs some auxiliary configuration. First of all, you should set its download directory with the following terminal command, provided that your download folder is called Downloads (be careful to use both ‘>’):
 … $ echo export XDG_DOWNLOAD_DIR=\"\$HOME/Downloads\" >> ~/.profile
This web browser is a very basic (but fast) one, which has the advantage of displaying web pages as GoldenDict would because it’s based on WebKit too, even if it has no buttons to navigate. It has just a few control keys which you can take a look at in your favorite browser with the following command:
 …$ x-www-browser http://uzbl.org/keybindings.php
Save that page from the browser to your download folder, and move it over to the configuration directory of uzbl (note the ’_’ replacing the space in the name):
 …$ mv ~/Downloads/Uzbl\ Keybindings.html ~/.config/uzbl/Uzbl_Keybindings.htm
This help file will be available in the Uzable Browser if it’s assigned to the [H] key (bear in mind that such keys are case sensitive in uzbl). This can be done by executing a last command, viz. the following one:
 …$ echo @cbind H = event REQ_NEW_WINDOW file://@config_home/uzbl/Uzbl_Keybindings.htm \
>> ~/.config/uzbl/config

Files tab  2. Dictionary files

Of the many dictionary packages available from the Software Manager, there are only a few which are really usable. So here’s the editor’s choice:
  • The Moby thesaurus is a large and comprehensive English thesaurus which concentrates on synonyms rather than on some fancy relationships between them as it is the case with thesauri derived from the WordNet project;

  • The bilingual dictionaries of the FreeDict project which can be found in the Software Manager by searching for “freedict”. They are quite good (and free), but the lack of a title once a package has been installed is somewhat embarrassing with regard to its handling in GoldenDict.
These local dictionaries have to be made known to GoldenDict, which is itself available as Menu› Education› GoldenDict. In order to configure it, you have to Edit› (its) Dictionaries…, or use the shortcut [F3] for this purpose. Under the [Files] tab, you should first Add… the following recursive path:
 /usr/share/dict
Then, create a new folder named .dict in your home folder. Naming it this way will make it invisible, unless you let Nemo View› Show Hidden Files (or toggle its display with [Ctrl+H]). Anyway, this will be the place where non-packaged dictionaries have to be downloaded, so state it as another recursive path in GoldenDict as:
 ~/.dict
Local dictionaries have to be indexed, which occurs automatically by quitting and restarting the program.

Dictionaries tab  3. Other dictionaries

If you need more than you can get from the Debian-Ubuntu-Linux Mint repositories, then you have two options basically, both of them being “free”, but in the two different meanings of the word:
  • First, there is a collection of dictionaries at http://www.personal.leeds.ac.uk/~ecl6tam/Downloadable/GoldenDict%20Dictionaries.html, compiled especially for GoldenDict by Alec McAllister from the University of Leeds. As the dictionaries therein all have GPL and analogous licenses, they’re free as in free speech 

  • Second, if you’re still not satisfied, you’ll have to swallow your pride, and take a look at the almost unique alternative that can be used. This one might be more professional, but it’s only as free as a free beer — at least at first sight. Because this beer is moldy: any dictionary on their website requires a valid license to use it 
In any event, the dictionaries have to be placed into the ~/.dict folder from above. It’s more clear to make up a subfolder there for each language, e.g. English, French, etc. If they come bundled within an archive of some sort, you have to unpack the dictionaries of your choice, for instance by dragging them from the Archive Manager to that folder. As mentioned in the precedent section, don’t forget to restart GoldenDict afterward.

Sound directories tab  4. Pronunciation

In online dictionaries, it’s a common feature nowadays to get the words pronounced. But to have that feature accessible offline, by means of audio files organized like a dictionary e.g., is even more challenging than to find a good dictionary. There are none in the repositories, and they are hard to find on the net.
Thus the only collection for English that I can recommend is the one that can be found on the website of the LightLang dictionary project which is of good quality and contains 21056 entries. Once downloaded, unpack the en folder of the archive into a (new) subdirectory of your dictionary folder which you should Add… to your dictionaries under the [Sound Dirs] tab. These sound resource has to be given a name too:
 Path: ~/.dict/sound/en
 Name: en Sound
 Icon: /usr/share/cinnamon/applets/keyboard@cinnamon.org/flags/gb.png
The local pronunciation files installed there can be listened to either by making GoldenDict: [F4] (Edit› Preferences›) Audio› Playback› Use (an) external program like mplayer, or by letting it Use (the) internal player, which is faster but needs espeak to be configured as an external (!) program.
Besides, there are also high quality packages for French which can be downloaded form the website of the Shtooka project. Unfortunately, they have a special format that needs a dedicated software to decode, which doesn’t compile easily on Linux Mint 17.x. Hence, it’s easier to use the site as it is, and to set it up as an online resource as shown in the next section.

Websites tab  5. Online resources

If we were to constitute a dictionary for English not as a foreign language, but as a native one, we could as well take French as the language to look up, so the Shtooka website could serve us as an example. That is, we should Add… it as a [Website] to our dictionaries with the following parameters:
 Enabled: ✕
 Name:::: Projet Shtooka
 Address: http://shtooka.net/listen/fra/%GDWORD%
The trick is to replace the word to look up, e.g. “dictionary” as in the screenshot above, by the template %GDWORD% in the address of the online dictionary’s query page. This works for any website that displays the targeted word in the address bar of a browser, so for example the dict.cc online dictionary which I consider one of the best resources for the language pair German–English. Hence, you could query this dictionary with the following URL:
  http://www.dict.cc/?s=dictionary
→ http://www.dict.cc/?s=%GDWORD%
Like any other resources, websites can be assigned icons in the [F3] Dictionaries› Sources› (or later the Groups›) dialog, provided you scroll it far enough to the right. For instance, you could choose an image from the folder /usr/share/cinnamon/applets/keyboard@cinnamon.org/flags. If you manage to download the site’s logo, you can use it too. A convenient place to store your icons would be another subdirectory of the dictionary folder (the icon’s location should be first copied from its context menu in Nemo, and then pasted into the dialog):
 ~/.dict/images
However, if a query web page doesn’t render in a GoldenDict window, e.g. because openjdk-7-jdk or any other Java package doesn’t allow for it (saying “Vector smash protection is enabled” instead), you’ll have to access it via an external program.

Programs tab  6. External commands

One of the most prominent websites that doesn’t work for me in GoldenDict is Google Translate. At the same time, it’s the most easiest thing in the world to make it work in a window of its own. All it takes is to Add… it as a [Program] with the following description:
 Enabled:::::: ✕
 Type::::::::: Html
 Name::::::::: Google 2 en
 Command Line: uzbl https://translate.google.com/?hl=en#auto/en/%GDWORD%
The “program” above is basically the same web address which would stand under [Websites], i.e. an address with the query word replaced by “%GDWORD%”, augmented with the name of the browser which is to display it. The command given will translate any fragment of text — by detecting its language automatically — to English which is specified as usual by en. If you need to translate to another target language, you have to substitute this with the code of that language in the description fields. Supposing your current session is running in this language, you can determine its code just by pasting the following command line into a terminal:
 …$ echo $LANG | head -c2 >> echo
The same principle applies to the word-reading application. As all these resources will be handled by GoldenDict in exact the same manner as any dictionary, a reader should be defined for each language you’re using. That is, the language code must match the target language, so for example in the case of English:
 Enabled:::::: ✕
 Type::::::::: Audio
 Name::::::::: en Speaker
 Command Line: espeak -v en %GDWORD%
 Icon::::::::: /usr/share/doc/espeak/docs/images/lips.png
Now, let’s suppose that you have downloaded some free grammar book to the dictionary folder as a PDF. You can then look up any grammatical term like in a dictionary simply by calling the Document Viewer, but you have to insert your username into the path leading to the file:
 Enabled:::::: ✕
 Type::::::::: Html
 Name::::::::: English Grammar
 Command Line:
evince --find=%GDWORD% file:///home/your_username_here/.dict/free-english-grammar.pdf
 Icon::::::::: /usr/share/icons/Mint-X/apps/scalable/evince.svg

Wikipedia tab  7. MediaWiki-based sites

GoldenDict comes with some Wikipedia and Wiktionary sites preconfigured, but I prefer to use the English Wiktionary exclusively for two reasons:
  • I’ve got Configurable Menu to look up things directly in the Wikipedia corresponding to my session already, and that works so well that I don’t need to duplicate its function;

  • From a linguistic standpoint, the English version of Wiktionary is the most complete and accurate, even with regard to entries in other languages, so it’s the most useful to me.
But if you don’t get along with this setting, you can configure any Wikipedia or Wiktionary site of your liking by adapting the default examples to your needs, of course.

Morphology tab  8. Spelling dictionaries

You don’t have to care much about these dictionaries, as they’re available by default in Linux Mint. That is to say, the spelling dictionaries of Hunspell get installed automatically when you add a new language in System Settings› Language (Settings)› Language› Language support, because it’s those packages that do the work in the background when LibreOffice or a web browser needs spell checking.
GoldenDict uses them to provide stem words and make spelling suggestions for your searches. Just mark the ones you need, and leave the directory path at its default setting which is:
 /usr/share/myspell/dicts




Groups tab  The right look-up order

Now that you’ve installed all your offline dictionaries and configured the other sources, you should organize them. Since anything you’ve added is regarded as a dictionary, you could set up the order that you wish the searches to follow under the tab [Dictionaries] simply by dragging them around. But as you’ll inevitably end up adding new dictionaries at the end of that list, arranging them there is superfluous.
Why ? Because your dictionaries should be grouped according to their language — which would be the language of your user interface ideally — and ordered following their importance (and speed). GoldenDict offers to sort all your dictionaries in language pairs, but I think this isn’t an idea as good as it sounds, since translations are often needed from and to the same language. Another important fact that impacts on grouping is the ability of a source to render within GoldenDict. As a starting point, you can take on the following advices drawn from my own experience with the matter.
With reference to our examples so far, you should Add (a) group named “English” in the [F3] Dictionaries› Groups› dialog by choosing “England” as a Group icon. You should add the following items from the Dictionaries available by dragging them into the Groups window — or by copying them by clicking on [>] — in the following order:
  1. a monolingual reference like the first dictionary in the screenshot above or the Merriam-Webster’s;

  2. a dictionary of synonyms like the Moby Thesaurus;

  3. your preferred bilingual dictionaries in the sequence of the frequency of their use (by taking care to place slower online resources down the list);

  4. specialized dictionaries like the Hacker’s Jargon or American English;

  5. the en Speaker, preceded by the en Sound, if available (these resources won’t be used directly unless the loudspeaker icon beneath the search field in GoldenDict doesn’t work for some reason);

  6. the English Wiktionary;

  7. and last but not least, the corresponding morphology dictionary.
If you have dictionaries that can only be displayed externally, you should put each of them in a group of its own. Following the path above, you should have three groups which should show England’s colors (the first to be defined in such a bunch of groups being the default):
  • “English” as set up like before;

  • “Google Translate” with Google 2 en as its sole dictionary;

  • “Grammar” with the English Grammar book.
Proceed like this for each language which you’ve configured dictionaries for, possibly assigning some of them to several groups, e.g. in the case of any bilingual variant. This way, the purpose of your groups will be clear, and they can be made to adapt to the desktop’s language.

Automatic language selection

This step isn’t imperative, it just makes it more convenient to return to the default dictionary group after having messed around with the others: you stop GoldenDict with [Ctrl+Q] from its main window or with the mouse in the panel, and the next time you need it again, it’ll start up with a dictionary group that matches the current language of your desktop.
It isn’t that difficult to install this feature though. It works via a small intermediary script that must be first made up with the following sequence of terminal commands (as usual you can copy’n paste them) — remember, the sudo command requires your password:
…$ sudo touch ~/usr/local/bin/goldendict
…$ sudo chmod +x ~/usr/local/bin/goldendict
…$ gedit ~/usr/local/bin/goldendict

The last command opens up the Text Editor where you’ll have to paste in the following content and to save the file — that’s it:
#!/bin/bash
# /usr/local/bin/goldendict (MagicMint) O0820
# A proxy for GoldenDict to set the correct dictionary group
# (ɔ) GPL-2, see /usr/share/common-licenses/GPL-2

CONF=~/.goldendict/config

# Take the first group which has a flag corresponding to the current language
GROUP=$(awk -F 'id="' "/group icon=\"`echo $LANG | head -c2`/{print \$2}" \
 $CONF | head -1 | cut -d'"' -f1)
echo Group id = $GROUP

# Make that group the default one
sed -i "/<lastMainGroupId>/s/<lastMainGroupId>.*</<lastMainGroupId>$GROUP</" $CONF
sed -i "/<lastPopupGroupId>/s/<lastPopupGroupId>.*</<lastPopupGroupId>$GROUP</" $CONF

# Start the viewer
exec /usr/bin/goldendict "$@"

#End of script




  The best of it all

As @xenopeek suggested it on the forums, single words can be looked up with a unique keypress from within any application, even from a terminal window or non-compliant editors like gvim. This idea is just extended here to an arbitrary selection of text which is much more useful in terms of translation. Note that for the following to work in GoldenDict, you must not alter the built-in hotkeys of the program !
All you have to do (in Cinnamon, e.g.) is to Add (a) custom shortcut in System Settings› Hardware› Keyboard› Keyboard shortcuts› Custom Shortcuts› with the following settings:
 Name:::: Dictionary
 Command:
bash -c "goldendict \"$(xsel|tr '\n' ' '|sed -r 's/^[^[:alpha:]]*([-[:alpha:]]*).*$/&/')\""
You should bind this command to the [Super (Windows) + D] key (unless you need that negligible “Show desktop” function on this key absolutely ). This makes it possible to look up things in your dictionaries very simply and quickly by following these steps:
  • Mark the text you want to translate by the mouse or the arrow keys;

  • Press [Super+D] to call GoldenDict. It will show its main window if it starts at the moment, or it will pop up a floating window if it was running already. From the latter, you can send the query by means of a click on the book icon or with [Ctrl+W] to the main window which has more possibilities, if you want to. However, if the program is running in the background, its window can be opened without any query at all by pressing [Ctrl+F11+F11] at any time;

  • In both windows, you can navigate either by clicking on a link (or the back and forward buttons), or by double-clicking on a word;

  • You can mark a synonym or a translation there and copy it with the mouse or [Ctrl+C] to the clipboard as usual. If the selection concerns a locution, you perform it by dragging the mouse pointer along it, otherwise it suffices to click on a single word to select it — provided you’ve got [F4] Preferences› Interface› Select word by single click ticked;

  • Once you’ve returned to your original application, you can paste the contents of the clipboard over the previously selected text in order to replace it — and you’re done in a jiffy… 
  • Source: https://community.linuxmint.com

Monday, February 27, 2017

Free/open-source machine translation software

Rule-based systems

  • Apertium, a free/open-source rule-based machine translation platform.
  • Matxin, a free/open-source rule-based machine translation system for Basque.
  • OpenLogos, a free/open-source version of the historical Logos machine translation system.
  • Anusaaraka, English-Hindi machine translation system.

Statistical machine translation systems

Decoders

  • Moses, a statistical machine translation system.
  • Marie, an n-gram-based statistical machine translation decoder.
  • Joshua, an open source decoder for statistical translation models based on synchronous context free grammars
  • Phramer, an open-source statistical phrase-based machine translation decoder
  • GREAT, a decoder based on stochastic finite-state transducers, which includes a training toolkit.
  • The Thot toolkit includes a decoder as of 2014.
  • Travatar is a tree-to-string statistical machine translation system.
  • CDEC is a decoder, aligner, and model optimizer for statistical machine translation and other structured prediction models based on (mostly) context-free formalisms written by Chris Dyer at the Language Technologies Institute in Carnegie Mellon University

Training translation models

  • Giza++ is a tool to train translation models for statistical machine translation (see also the related mkcls tool to train word classes)
  • Thot includes a toolkit to train phrase-based models for statistical machine translation.

Language models

  • IRSTLM, free/open-source language modelling tool to be used with Moses instead of SRILM, which is not free.
  • RandLM, space-efficient ngram-based language models built using randomized representations (Bloom Filters etc).
  • Kenneth Heafield's software for the fast filtering of ARPA format language models to multiple vocabularies.
  • Holger Schwenk's Continuous Space Language Model toolkit (CSLM) works by projecting the word indices onto a continuous space and using a probability estimator operating on this space.

Scoring

  • Kenneth Heafield's scripts that make it easy to score machine translation output using NIST's BLEU and NIST, TER, and METEOR.

Other software

  • RIA is a tool for automatic induction of transfer rules for Transfer-Based Statistical Machine Translation using dependency structures.
  • Chaski: Distributed phrase-based machine translation training tool based on Hadoop.
  • Grammatical Framework, a free/open-source programming language used to create grammars for multilingual applications.

Example-based machine translation systems

Multi-engine machine translation / system combination

Aligners and translation models

  • Giza++: training of statistical translation models.
  • Anymalign, a multilingual sub-sentential aligner.
  • Ventsislav Zhechev's Sub-tree aligner which can be used for the automatic generation of parallel treebanks.

Web services around machine translation

  • Tradubi is an open-source Ajax-based web application for social translation built upon Apertium (may be tested online).

Distributed machine translation

  • ScaleMT (no release yet, browse at the Apertium Subversion repository) is a free/open-source framework for building scalable machine translation web services.

Quality estimation

  • Quest++, an open source tool for translation quality estimation developed by the group of Lucia Specia at the Univ. of Sheffield (note that the current version still has one important non-free dependency: SRILM).

Other useful tools

... that may be used to build machine translation systems
  • Freeling, a free/open-source suite of language analyzers.
  • Bitextor, an automatic bitext harvester
  • Foma, a finite-state machine toolkit and library
  • HFST, Helsinki Finite State Technology for natural-language morphologies.
  • VISL CG-3, the constraint grammar parser at the Visual Interactive Syntax Learning project of Syddansk Universitet: browse Subversion repository, source snapshots.
  • Source: http://fosmt.org/
Other tools, some may be outdated (source: www.word2word.com)