Lab 6: Similarity-Based Prediction on Tabular Data
CS1234: Small-Scale Application Development


Problem Description:

You are given a training table containing several records.
Each record contains multiple attributes followed by a final label.

A row in the table has the form:

x1 x2 x3 ... xn y


Where:
x1 ... xn → attribute values (features)
y → label (prediction output)



You are also given a query record:
q1 q2 q3 ... qn


Your task is to predict the label y for this query.

** Prediction Rule **
-> To determine the prediction:

* Compare the query with each row in the table.
* Count how many attributes match exactly.
* The row with the maximum number of matching attributes determines the prediction.
* Output the label y of that row.



** Tie-Breaking Rule**

-> If multiple rows obtain the same highest number of matches:
* The label of the topmost (earliest) row in the table must be chosen.

** No-Match Case **

* If every row has zero matching attributes, output:
-1

** Input Format **

Each test case consists of two files:

1) Table File

tb.txt

Contains multiple rows.
Each row contains space-separated strings:

x1 x2 x3 ... xn y


All rows will have the same number of attributes.

The last column is always the label.

2) Query File

q.txt

Contains a single line:

q1 q2 q3 ... qn

The number of query attributes will match the number of attributes in the table rows.


Your script will be executed as:
awk -f solution.awk input/input01/q.txt input/input01/tb.txt



** Output Format **

Print a single line:

<label>


OR

-1

Constraints

-> 1 ≤ number of rows in table ≤ 5000
-> 1 ≤ number of attributes (n) ≤ 200
-> Each attribute is a non-empty string without spaces
-> Labels are also strings
-> All comparisons are case-sensitive
-> Query attribute count will always match the table attribute count

Sample Test Case 0:
 
tb.txt
red small round apple
yellow long curved banana
green small round guava
red round big watermelon

q.txt
red small round

Output
apple

Explanation :

We compare the query with every row:

Row	Matching Attributes
red small round apple	3
yellow long curved banana	0
green small round guava	2
red round big watermelon	1

The first row has the maximum number of matches → label apple.

Sample Test Case 1 (Tie Case) :

tb.txt
A B C X
A B D Y
A E C Z

q.txt
A B Q

Output
X

Explanation:

Row	Matches
A B C X	2
A B D Y	2
A E C Z	1

Two rows tie with 2 matches.
The topmost row is selected → X.

Sample Test Case 2 (No Match):

tb.txt
cat black small pet
dog brown big pet
cow white large farm

q.txt
car blue fast

Output
-1

Explanation:

No row shares any matching attribute with the query.

** Notes & Submission Guidelines **

* Your solution must work correctly for multiple test cases.
* The program must read input strictly from the provided files.
* No manual modification of test files is allowed.
* Output must exactly match the required format (no extra spaces or lines).

** Follow-Up Questions (Ungraded Practice) **

1) Instead of selecting only one best row, output the labels of the top 3 most similar rows.
2) What if Query's attribute count <= Table's attribute count ?
3) Modify the system so that earlier attributes are more important than later attributes.(e.g., x1 is more important than x2).
4) Instead of one query, process a file containing multiple query rows and print predictions for all of them.
