# Open file in folder and subfolders

**URL:** <https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552>\
**Category:** Coding Questions\
**Created:** [February 4, 2020, 6:08pm UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552 "2020-02-04T18:08:47Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![plux](https://avatars.discourse-cdn.com/v4/letter/p/91b2a8/32.png) [@plux](https://discourse.processing.org/u/plux)\
**Post date:** [February 4, 2020, 6:08pm UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/1 "2020-02-04T18:08:47Z")

</div>

Hello,  
I had this code working at some point. However I’m not sure what I’ve changed and now it’s broken.  
The method should open every file in a folder and in each subfolders (recursively) to check which one has less lines.  
Assuming all files/folders are inside data folder, I can access individually each file as

```auto
// working
String[] s = loadStrings("/subfolder/test.csv");
println(s.length);

```

but the following does not work (anymore).  
the structure I have is

```auto
/data
  /subfolder
    /test.csv
  /subfolder2
    /test2.csv

```

```auto
File folder;

void setup() {  
  folder = new File(dataPath(""));
  println(folder); //prints correctly

  int lineCount = 100000;
  lineCount = lineCounter(folder, lineCount);
}

int lineCounter(File folder, int lineCount) {
  for (File fileEntry : folder.listFiles()) {
      if (fileEntry.isDirectory()) {
          lineCounter(fileEntry, lineCount);
      } else {
          System.out.println(fileEntry.getName());
          String[] lines = loadStrings(fileEntry.getName()); // prints filename correctly
          if(lines.length < lineCount) // FAILS with The file "test.csv" is missing or inaccessible....
            lineCount = lines.length;
      }
  }
  return lineCount;
}

```

---

<div class="post-metadata">

**Author:** ![GoToLoop](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/gotoloop/32/86_2.png) [@GoToLoop](https://discourse.processing.org/u/GoToLoop)\
**Post date:** [February 4, 2020, 7:38pm UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/2 "2020-02-04T19:38:58Z")

</div>

> [@plux](#):
>
> … open every file in a folder and in each subfolders (recursively)…

You can use **listPaths()** for it:  
[Processing.GitHub.io/processing-javadocs/core/processing/core/PApplet.html#listPaths-java.lang.String-java.lang.String…-](http://processing.github.io/processing-javadocs/core/processing/core/PApplet.html#listPaths-java.lang.String-java.lang.String...-)

And here’s an example for image files, which you can easily adapt for CSV 1s:

> [@File Listing Order](https://discourse.processing.org/t/file-listing-order/7148/41):
>
> [Docs.Oracle.com/en/java/javase/11/docs/api/java.base/java/util/Arrays.html#sort(java.lang.Object[])](http://Docs.Oracle.com/en/java/javase/11/docs/api/java.base/java/util/Arrays.html#sort(java.lang.Object%5B%5D)) static final String PICS\_EXTS = "extensions=,png,jpg,jpeg,gif,tif,tiff,tga,bmp,wbmp"; static final float FPS = .25; PImage[] images; void setup() { size(1200, 600); frameRate(FPS); imageMode(CENTER); final File dir = dataFile(""); println(dir); String[] imagePaths = {}; int imagesFound = 0; if (dir.isDirectory()) { imagePaths = listPaths(dir.getPath(), "files", "recursi…

Don’t forget Processing can also load CSV & TSV files via **loadTable()**:

> **[loadTable() / Reference](https://processing.org/reference/loadTable_.html)**
>
> Reads the contents of a file or URL and creates a Table object with its values. If a file is specified, it must be located in the sketch's "data" folder. The filename parameter can also be a URL to a …

---

<div class="post-metadata">

**Author:** ![plux](https://avatars.discourse-cdn.com/v4/letter/p/91b2a8/32.png) [@plux](https://discourse.processing.org/u/plux)\
**Post date:** [February 5, 2020, 4:53pm UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/3 "2020-02-05T16:53:07Z")

</div>

Hello, thanks for the links!  
What I have to do is open about 400 .csv files (each file has about 5000 rows), determine the one with the least number of rows and then trim all the other files (deleting from the last row) so that they all have the same number of rows. What do you think is the best way to do it? LoadStrings and BufferWriter or using the Table class?  
Thanks

---

<div class="post-metadata">

**Author:** ![GoToLoop](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/gotoloop/32/86_2.png) [@GoToLoop](https://discourse.processing.org/u/GoToLoop)\
**Post date:** [February 5, 2020, 5:43pm UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/4 "2020-02-05T17:43:50Z")

</div>

> [@plux](#):
>
> **loadStrings()** and BufferedWriter or using the Table class?

Well, **loadTable()** automatically parses CSV & TSV files for us:

> **[loadTable() / Reference](https://processing.org/reference/loadTable_.html)**
>
> Reads the contents of a file or URL and creates a Table object with its values. If a file is specified, it must be located in the sketch's "data" folder. The filename parameter can also be a URL to a …

And by using the Table class we have access to useful methods such as **getRowCount()**:

> **[getRowCount() / Reference](https://processing.org/reference/Table_getRowCount_.html)**
>
> Returns the total number of rows in a Table.

So you can find out the Table w/ the least number of rows.

Then afterwards, invoke **setRowCount()** on all 400+ Table containers, so they’ll all have the same number of rows:  
[Processing.GitHub.io/processing-javadocs/core/processing/data/Table.html#setRowCount-int-](http://Processing.GitHub.io/processing-javadocs/core/processing/data/Table.html#setRowCount-int-)

And finally, use **saveTable()** on each trimmed down Table:

> **[saveTable() / Reference](https://processing.org/reference/saveTable_.html)**
>
> Writes the contents of a Table object to a file. By default, this file is saved to the sketch's folder. This folder is opened by selecting "Show Sketch Folder" from the "Sketch" menu. Alter…

---

<div class="post-metadata">

**Author:** ![plux](https://avatars.discourse-cdn.com/v4/letter/p/91b2a8/32.png) [@plux](https://discourse.processing.org/u/plux)\
**Post date:** [February 6, 2020, 11:13am UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/5 "2020-02-06T11:13:47Z")

</div>

Ok I still had to create File objects because I have to name each file depending on the parent folder, but this seems to be working well. Don’t know if there’s an easier solution for the file naming. Thanks!

```auto
void setup() {  
  String[] files = listPaths(dataPath(""), "files", "recursive", "extensions=,csv");
  int rowCount = 100000;
    
  for(int i = 0; i < files.length; i++){
    Table table = loadTable(files[i]);
    if(table.getRowCount() < rowCount)
      rowCount = table.getRowCount();
  }
  
  int index = 1;
  String parent = "";
  for(int i = 0; i < files.length; i++){
    if(!parent.equals(new File(files[i]).getParent()))
      index = 1;

    parent = new File(files[i]).getParent();  
    Table table = loadTable(files[i]);
    table.setRowCount(rowCount - 1);
    saveTable(table, parent + "_p" + (index++) + ".csv");
  }
}

```

---

<div class="post-metadata">

**Author:** ![GoToLoop](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/gotoloop/32/86_2.png) [@GoToLoop](https://discourse.processing.org/u/GoToLoop)\
**Post date:** [February 7, 2020, 3:51am UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/6 "2020-02-07T03:51:01Z")

</div>

> [@plux](#):
>
> Ok I still had to create File objects…

You can replace **listPaths()** w/ **listFiles()** in order to get a File[] instead of a String[] array:  
[Processing.GitHub.io/processing-javadocs/core/processing/core/PApplet.html#listFiles-java.io.File-java.lang.String…-](http://processing.github.io/processing-javadocs/core/processing/core/PApplet.html#listFiles-java.io.File-java.lang.String...-)

BtW, I see you use **loadTable()** in 2 places! That makes the code doubly slower!  
You should store the 1st batch in an array, so you can re-read from it later.

> [@plux](#):
>
> Don’t know if there’s an easier solution for the file naming.

Not easier, but I’ve come up w/ an alternative solution which is safe-proof against cases where **listFiles()** may return folders outta order, thus messing up w/ your index-renaming scheme.

For that I’m relying on a HashMap of ArrayList of Table containers mapped to a String key representing their parent folder name:

- [HashMap / Reference / Processing.org](http://Processing.org/reference/HashMap.html)
- [ArrayList / Reference / Processing.org](http://Processing.org/reference/ArrayList.html)

So all files belonging to 1 subfolder is assured to be processed in 1 batch only.

> [@plux](#):
>
> … because I have to name each file depending on the parent folder,

I couldn’t figure out exactly whether those renamed file names include their parent folder name as well or it’s only the index value as name.

But in my version here, I’m just dumping all renamed CSV files in a subfolder named “output/”, leaving the original 1s intact:

```auto
// Discourse.Processing.org/t/open-file-in-folder-and-subfolders/17552/6
// GoToLoop (2020-Feb-07)

import java.util.Map;
import java.util.List;

static final String SEARCH = "extensions=,csv", DST = "output/", CSV = ".csv";

void setup() {
  final File root = dataFile("");
  println(root);
  if (!root.isDirectory()) System.err.println("/data subfolder not found!");

  final File[] files = listFiles(root, "files", "recursive", SEARCH);
  final int len = files.length;
  println(len, "csv tables found.");

  final Map<String, List<Table>> tableDirs = new HashMap<String, List<Table>>();
  int minRow = MAX_INT;

  for (final File f : files) {
    final String dir = f.getParentFile().getName();

    List<Table> tables = tableDirs.get(dir);
    if (tables == null) tableDirs.put(dir, tables = new ArrayList<Table>());

    final Table t = loadTable(f.getPath(), "header");
    tables.add(t);
    minRow = min(minRow, t.getRowCount());
  }

  println("Subfolders w/ CSV files inside:", tableDirs.keySet());
  if (len > 0) println("Smallest loaded table had", minRow, "row(s).");

  for (final Map.Entry<String, List<Table>> entry : tableDirs.entrySet()) {
    final String dir = entry.getKey();
    int idx = 1;

    for (final Table t : entry.getValue()) {
      t.setRowCount(minRow);
      if (t.getColumnCount() >= 3) for (int i = 0; i++ < 3; t.removeColumn(0));
      saveTable(t, DST + dir + "_p" + idx++ + CSV);
    }
  }

  exit();
}

```

---

<div class="post-metadata">

**Author:** ![plux](https://avatars.discourse-cdn.com/v4/letter/p/91b2a8/32.png) [@plux](https://discourse.processing.org/u/plux)\
**Post date:** [February 11, 2020, 2:12pm UTC](https://discourse.processing.org/t/open-file-in-folder-and-subfolders/17552/7 "2020-02-11T14:12:19Z")

</div>

wow, I’ll give it a go as soon as I have time, it looks great! Didn’t know about loadFile()!  
Thanks a lot!
