WN_CONNECT source v1.1 : 
%%%%%%%%%%%%%%%%%%%%%%%%

WN_CONNECT is a software application prototype that aims to access the lexical database WordNet. One of its main features is that it has been fully implemented using PROLOG or related technologies.

WN_CONNECT is divided into the following modules.

:- use_module(wn).
:- use_module(wn_synsets).
:- use_module(wn_hypernyms).
:- use_module(wn_similar_adjectives).
:- use_module(wn_sim_measures).
:- use_module(wn_ic_measures).
:- use_module(wn_rel_measures).
:- use_module(wn_display_graph).
:- use_module(wn_gen_prox_equations).

A general characteristic of the predicates implemented in these modules is that the parameter Word (occurring in that predicates) is a term that follows the syntax "Word[:SS_type[:Sense_num]]". 

Where SS_type is a one character code indicating the synset type:
    n NOUN
    v VERB
    a ADJECTIVE
    s ADJECTIVE SATELLITE
    r ADVERB

and "Sense_num" specifies the sense number (meaning) of the word, within the part of speech encoded in the synset_id. "Sense_num" is a natural number: 1, 2, 3, ...

Note that sometimes this term may be partially specified; that is, SS_type and Sense_num could be variables (or even omitted).


%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn).
%%%%%%%%%%%%%%%%%%%%%%%%
This module was implemented by Jan Wielemaker. It discloses the Wordnet Prolog files in a more SWI-Prolog friendly manner. It exploits SWI-Prolog demand-loading  and SWI-Prolog Quick Load Files to load `just-in-time' and as quickly as possible.

The system creates Quick Load Files for  each wordnet file needed if the .qlf file doesn't exist and  the   wordnet  directory  is writeable. For shared installations it is adviced to   run  load_wordnet/0 as user with sufficient privileges to create the Quick Load Files.


%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_synsets).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements predicates to retrieve information about words and synsets stored in WordNet.

This module uses:

:- if(getenv('WNDEVEL', yes)).
:- use_module(wn_portray).
:- else.
:- use_module(wn).
:- endif.
 

The public predicates implemented in this module are:

    wn_word_info/1, 
    wn_gloss_of/2,
    wn_synset_ID_of/2,
    wn_synset_of/2,
    wn_synset_components/2,
    wn_synset_components/3
 

** wn_word_info(+Word): 
Prints the information about a word, Word, stored in the
database 'wn_s.pl' of Wordnet. 

** wn_gloss_of(+Word, -Gloss): 
Returns the Gloss of a Word. 

** wn_synset_ID_of(+Word, -W_Synset_ID): 
W_Synset_ID is the synset ID to which W belongs. 

** wn_synset_of(+Word, -W_synset): 
W_synset is the synset to which Word belongs. W_synset is represented by a set of words that are synonyms of Word.

** wn_synset_components(+Synset_ID, -Synset_Words):
Synset_Words is the list of words that compounds the synset Synset_ID.
It is equivalent to calling wn_synset_components/3 setting the third parameter to ``verbose(no)''. That is,
			wn_synset_components(Synset_ID, Synset_Words, verbose(no)).


** wn_synset_components(+Synset_ID, -Synset_Words, Verbosity):
Synset_Words is the list of words that compounds the synset Synset_ID.

"Verbosity" is a parameter that controls the degree of information shown. It can be set to "verbose(no)" or "verbose(yes)". In the second case for each word W in the synset, ist syntactic type SS_type and its sense number SS_Num are shown.



%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_hypernyms).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements predicates to retrieve information about hypernyms of a concept (synset). These predicates only work with either nouns or verbs. 

This module uses:

:- use_module(wn_synsets).
:- use_module(wn_display_graph).
 

The public predicates implemented in this module are:

	wn_hypernyms/2,
     wn_display_hypernyms/1,
     wn_display_graph_hypernyms/1,
     wn_hypernyms/3 (NOT DESCRIVED HERE. ONLY FOR INTERNAL USE: 
                     HYPERNYMS CHAINS FOR SIMILARITY MEASURES)


** wn_hypernyms(+Word, -List_SynSet_HyperNym):
Given a word (term) "Word" returns the list "List_SynSet_HyperNym" of its hypernym synset_IDs. Word is a hyponym of the elements of that list.

** wn_display_hypernyms(+ Word):
Given a word (term) "W_Hyponym" prints a textual representation of the chain of its hypernym synsets. Word is a hyponym of the elements of that chain.

** wn_display_graph_hypernyms(+Word)
It shows a graphic representation of all hypernyms corresponding to all the senses of the word 'Word'. A node of the graph only shows the representative word of that hypernym synset.



%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_similar_adjectives).
%%%%%%%%%%%%%%%%%%%%%%%%
This module defines predicates that find the adjectives which are similar
in meaning to a input adjective. Do not confuse "similar" with "synonym".
Synonym words are grouped in a synset and they are words equals in meaning.
In other words, synonym words must have the maximum (top) degree of similarity.

Most of predicates defined in this module a based on the operator "sim"

-------------------------
sim(synset_id,synset_id). (THIS MODULE IS IN CONSTRUCTION YET)
-------------------------
The "sim" operator specifies that the second synset is similar in meaning
to the first synset. This means that the second synset is a satellite of
the first synset, which is the cluster head. This relation only holds for
adjective synsets contained in adjective clusters.

The two addressed synsets are either two head synsets, or one head synset
and one satellite synset. There is no matching sim clause for two satellite
synsets. Because, if they would have similar meanings, they would be grouped
together in one synset.

Therefore, the predicates defined in this module are for adjectives and do not
work for other parts of speech.

This module uses:

:- use_module(wn_synsets).


The public predicates implemented in this module are:

	wn_sim_adjectives_of/2,
	wn_display_sim_adjectives_of/1


** wn_sim_adjectives_of(Word, List_sim_SynSets): 
It is true if List_sim_SynSets is a list of similar adjectives synsets of the adjective Word.
Note that Word must be an adjective. Only adjectives can be similar one of each other. That is, words of type "a" or "s". There is no matching sim clause for two satellite synsets. Because, if they would have similar meanings, they would be grouped together in one synset.

** wn_display_sim_adjectives_of(+Adjective):
Given a word (term) "Adjective", prints the list of its similar synsets.




%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_sim_measures).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements predicates to compute standard similarity measures between concepts based on counting edges.  

This module uses:

:- use_module(wn_hypernyms).
:- use_module(wn_synsets).
 

The public predicates implemented in this module are:

	wn_lcs/4,
	wn_path/3,
	wn_path_nondet/3,
	wn_wup/3,
	wn_wup_nondet/3,
	wn_lch/3,
	wn_lch_nondet/3


** wn_path(+Word1:SS_type1:W1_Sense_num, +Word2:SS_type2:W2_Sense_num, -Degree):
This predicate implements the PATH similarity measure.
Takes two concepts (terms -- Word:SS_type:Sense_num) and returns the degree of similarity between them. Note that we do not explicitly require information about the synset type and sense number of a word (that can be variables).

We check that both Word1 and Word2 are nouns or verbs but not combinations of them.

	sim_PATH(c1, c2) = 1/len(c1, c2)

	NOTE: len(W1, W2) = (DepthW1-LCS_depth) + (DepthW2-LCS_depth) +1

** wn_path_nondet(+Word1:SS_type:W1_Sense_num, +Word2:SS_type:W2_Sense_num, -Degree):
Nondeterministic predicate. It is the user interface to the local predicate path/3.
 
Inspects a pair of HyperTrees associated to Word1 and Word2 and obtains the degree of similarity between Word1 and Word2 (according to that pair of HyperTrees).

** wn_wup(+Word1, +Word2, -Degree)
This predicate implements the WUP similarity measure.
Takes two concepts (terms -- Word:SS_type:Sense_num) and returns the degree of similarity between them. Note that we do not explicitly require information about the synset type and sense number of a word (that can be variables).

We check that both Word1 and Word2 are nouns or verbs but not combinations of them.

	sim_WUP(c1,c2)= 2*depth(lcs(c1,c2)) / (Depth(c1)+Depth(c2))

** wn_wup_nondet(+Word1:SS_type:W1_Sense_num, +Word2:SS_type:W2_Sense_num, -Degree):
Nondeterministic predicate. It is the user interface to the local predicate wup/3.
 
Inspects a pair of HyperTrees associated to Word1 and Word2 and obtains the degree of similarity between Word1 and Word2 (according to that pair of HyperTrees)

** wn_lch(+Word1, +Word2, -Degree)
This predicate implements the LCH similarity measure.
Takes two concepts (terms -- Word:SS_type:Sense_num) and returns the degree of similarity between them. Note that we do not explicitly require information about the synset type and sense number of a word  (that can be variables).

We check that both Word1 and Word2 are nouns or verbs but not combinations of them.

	sim_LCH (c1, c2) = −ln[ len(c1,c2) / (2 * max{depth(c)|c in WordNet})]

	NOTE 1: len(W1, W2) = (DepthW1-LCS_depth) + (DepthW2-LCS_depth) +1
	NOTE 2: max{depth(c)|c in WordNet} is the maximum depth of a concept in 
		   the WordNet data base. In practice, is a fixed constant for each 
		   part of speech
				MaxDepth(n) = 20    (Nouns)
				MaxDepth(v) = 14    (Verbs)

** wn_lch_nondet(+Word1:SS_type:W1_Sense_num, +Word2:SS_type:W2_Sense_num, -Degree):
Nondeterministic predicate. It is the user interface to the local predicate lch/3.

Inspects a pair of HyperTrees associated to Word1 and Word2 and obtains the degree of similarity between Word1 and Word2 (according to that pair of HyperTrees).


FOR COMPUTING THE LESS COMMON SUBSUMER (LCS)
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
** wn_lcs(+Word1, +Word2, -LCS, -LCS_depth)
Returns the Less Common Subsumer LCS of two words and its depth LCS_depth from the root of the HyperTree

	NOTE:
	This predicate is nondeterministic.
	It can be used without specifying the Type and Sense of a Word, but in this case the LCS for all combinations of types and senses of these two words are obtained.

	If you want to obtain the LCS for two precise concepts introduce 
			Word1:W1_Type:W1_Sense
	and 
			Word2:W2_Type:W2_Sense



%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_ic_measures).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements predicates to compute standard similarity measures between concepts based on information content (IC).  

This module uses:

:- use_module(wn_hypernyms).
:- use_module(wn_synsets).
 

The public predicates implemented in this module are:

        wn_res/3,
        wn_jcn/3,
        wn_lin/3,
        information_content/3,
        frequency_of_use/2,
        gen_all_hyponyms_of/2


	NOTE: In all these IC measures
		 "Word1:SS_type1:W1_Sense_num" denotes de concept c1 and 
		 "Word2:SS_type2:W2_Sense_num" de concept c2.

** wn_res(+Word1, +Word2, -Degree):
This predicate implements the RESNIK similarity measure, based on information content.
Takes two concepts (terms -- Word:SS_type:Sense_num) and returns the degree of similarity between them. Note that we do not explicitly require information about the synset type and sense number of a word. 

We check that both Word1 and Word2 are nouns or verbs but not combinations of them.
		sim_RES(c1,c2)= IC(lcs(c1,c2))


** wn_jcn(+Word1, +Word2, -Degree):
This predicate implements the JIANG & CONRATH similarity measure, based on information content.
Takes two concepts (terms -- Word:SS_type:Sense_num) and returns the degree of similarity between them. Note that we do not explicitly require information about the synset type and sense number of a word. 

We check that both Word1 and Word2 are nouns or verbs but not combinations of them.
		sim_JCN(c1,c2)= 1/ [IC(c1) + IC(c2) - 2*IC(lcs(c1,c2))]


** wn_lin(+Word1, +Word2, -Degree):
This predicate implements the JIANG & CONRATH similarity measure, based on information content.
Takes two concepts (terms -- Word:SS_type:Sense_num) and returns the degree of similarity between them. Note that we do not explicitly require information about the synset type and sense number of a word. 

We check that both Word1 and Word2 are nouns or verbs but not combinations of them.
		sim_JCN(c1,c2)= 1/ [IC(c1) + IC(c2) - 2*IC(lcs(c1,c2))]


** information_content(+Synset_ID, +Root_ID, -IC):
Computes the information content IC of the concept denoted by the Synset_ID. This
quantity is defined as:

		IC = -ln(Frequency/Frequency_Root)

if frequency_of_use(Synset_ID, Frequency) and frequency_of_use(Root_ID, Frequency_Root)

	NOTES:
	1) Root_ID is the synset number of the concept in the root of the hierarchy.
	2) IC(c) is defined as the natural logarithm of the probability of 
	   encountering an instance of a concept c (meassured in terms of a relative 
	   frequency of use of the concept c in a corpus).
	3) Natural logarithm: logarithm to the base of the mathematical constant e.


** frequency_of_use(+Synset_ID, -Frequency),
In WordNet a synset represents a concept (which is identified by a Synset_ID). This predicate computes the frequency of use of a concept in a corpus (using the information stored in the last parameter of the predicate wn_s/6). We compute the "synset tag sum" of the concepts subsumed by Synset_ID adding them to obtain the Frequency of the concept Synset_ID.

Note that the concepts subsumed by Synset_ID are their hyponyms


** gen_all_hyponyms_of(+Synset_ID, -List_all_Hyponym_IDs).
"List_all_Hyponym_IDs" is the list of all hyponyms of the synset "Synset_ID". The list "List_all_Hyponym_IDs" is a bag of synset_IDs hyponyms of "Synset_ID".

	NOTES: Because WordNet hierarchies are not actually trees, some synset 
		nodes are present several times. This can lead to overestimate some 
		word frequencies of use.


%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_rel_measures).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements a new relatedness measures between concepts based on the Jaccard index. THIS MEASURE HAS NOT A GOOD PERFORMANCE. IT HAS TO BE IMPROVED.

This module uses:

:- use_module(etu).               %%% Michael A. Covington's Efficient Tokenizer
:- use_module(wn_synsets).
:- use_module(library(snowball)). %%% The Snowball multi-lingual stemmer library
 

The public predicates implemented in this module are:

	wn_yarm/3

** wn_yarm(+Word1, +Word2, -Degree)
YARM (Yet Another Relatedness Measure) compares the gloses SGL_W1 and SGL_W2 of two words (after removing stop words and stemming) by computing the Jaccard index for them as the relatednes Degree:
  Degree = |SGL_W1 intersect SGL_W2| / |SGL_W1 union SGL_W2|
		= |SGL_W1 intersect SGL_W2| / (|SGL_W1| + |SGL_W2| - |SGL_W1 intersect SGL_W2|)



%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_display_graph).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements display_graph/1 for graphical display of graphs. It is an auxiliary module.

This module uses:

NONE 

The public predicates implemented in this module are:

	display_graph/1


** display_graph(+Graph)
	Graph is a list of arc(From,To)
	Displays a PDF containing the graphical representation of Graph
	Creates the files:
	- out.dot: A file with the graph in DOT format (graph description language) 
	- out.pdf: The PDF document with the graph representation
	- out.tex: The LaTeX document with the graph representation. Disabled for now (just uncomment it below for enabling)

% Use:
	display_graph(Graph), where Graph=[arc(v1,v2),...,arc(vn-1,vn)]

% Examples:
	?- findall(arc(X,Y),(wn_hypernyms(man,List),append(_,[X,Y|_],List)),Graph), display_graph(Graph).
	?- setof(arc(X,Y),List^H^T^(wn_hypernyms(man,List),append(H,[X,Y|T],List)),Graph), display_graph(Graph).

% Requires:
	- PDF displayer (as indicated in pdf_displayer/1 fact and accesible in the path).
	- dot (part of Graphviz, accesible in the path)
	- dot2tex (for generating a LaTeX version of the graph). If LaTeX output is enabled (disabled by default)


%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
:- use_module(wn_gen_prox_equations).
%%%%%%%%%%%%%%%%%%%%%%%%
This module implements predicates that help to generate proximity equations semi-automatically using WordNet information of the relatedness of concepts.

This module uses:

:- use_module(wn_sim_measures).
:- use_module(wn_utilities).
:- use_module(utilities).
 

The public predicates implemented in this module are:

  wn_gen_ontology_file/3,       % +ListOfListOfWords, +File, +Measure
  wn_gen_prox_equations_list/3, % +ListOfListOfWords, +Measure, -Equations
  wn_auto_gen_prox_equations/4  % +Directives, +Rules, -InEquations, -OutEquations


** wn_gen_ontology_file(+ListOfListOfWords, +File, +Measure):
Given a list of list of words, ListOfListOfWords, the name of a file, File, and the acronym of a measure, Measure (by now [path, wup, lch, res, lin, jcn, yar]), it generates a set of proximity equations and stored them into the file File.

** wn_gen_prox_equations_list(+ListOfListOfWords, +Measure, -Equations):
Given a ListOfListOfWords computes all proximity equations that can be formed paring the words of each list between them and then computing their proximity degree using the measure Measure.

	NOTES: Each list of ListOfListOfWords must be compounded by words of the 
		same part of speech (either nouns, verbs or adjectives)

		"sim(Word1, Word2, Degree)" is the internal Bousi~Prolog representation 
		of a proximity equation "Word1 ~ Word2 = Degree" (i.e., Word1 is close 
		to Word2 with approximation degree Degree).


** wn_auto_gen_prox_equations(+Directives, +Rules, -InEquations, -OutEquations):
If Directives is [:- directive(wn_gen_prox_equations, [Measure, Auto])], then return in OutEquations all the equations derived from constants in Rules. Otherwise, return InEquations

	NOTE: Use the following in setof if predicate names are required,
		and extract functors from the term (atoms_in_term):
		utilities:remove_program_prefix(Atom, Word), 
