3

I found myself wanting to do this in Elixir:

re_sentence_frag = %r/(\w([^\.]|\.(?!\s|$))*)(?=\.(\s|$))/
Regex.replace(re_sentence_frag, " oh.  a DOG. woOf. ", String.capitalize("\\1"))

Of course, that has no effect. (It capitalizes the string "\\1" just once.) What I really meant is to apply String.capitalize/1 to every match found by the replace function. But the 3rd parameter can't take a function reference, so passing &(String.capitalize("\\1") also doesn't work.

This seems so fundamental that I'm surprised it's not possible. Is there another approach that would as neatly express this kind of manipulation? It looks like the underlying Erlang libraries would not immediately support passing a function reference as the 3rd parameter, so this may not be completely trivial to fix in Elixir.

How would you program manipulation of each matched string?

7
  • The "\\1" is meant for consumption by the regex engine, not the String class. Commented Jan 22, 2014 at 17:54
  • I would look if a function ref is a parameter option. Where the function recieves match results and returns the replacement string. If it can't do that, then you have to reconstruct a new string in a regex find loop. Commented Jan 22, 2014 at 17:57
  • Your best bet is use scan and use the information from the result to manually replace them. For code-reuse purpose, you can create a wrapper function that accepts a function as parameter. Commented Jan 22, 2014 at 17:58
  • @nhahtdh, I think you're correct, although I was leaning toward using split. One of the goals above was to avoid changing the in-between bits, and it's not clear how to get those (or include them in the final results) using scan. I'll post one possible answer based on split. Commented Jan 22, 2014 at 21:36
  • You are right. We can't pass a function to the Erlang side so it is non trivial to support this feature. :( Split seems to be the best way to go for this case. Commented Jan 22, 2014 at 23:34

1 Answer 1

2

Here is one solution based on split:

" oh.  a DOG. woOf. pi is 3.14159. try version 7.a." |>
String.split(%r/(^|\.)(\s+|$)/)                      |>
Enum.map_join(&String.capitalize/1)

I guess it's not much more clumsy than my original attempt. The regex is considerably simpler, as it only needs to find the bits between sentences.

Sign up to request clarification or add additional context in comments.

1 Comment

I left some comments on split under your original posted question. Don't know if it will help or not.

Your Answer

By clicking “Post Your Answer”, you agree to our terms of service and acknowledge you have read our privacy policy.

Start asking to get answers

Find the answer to your question by asking.

Ask question

Explore related questions

See similar questions with these tags.